Content auditing method and device, electronic equipment and medium
By processing multimodal description data on product information pages of e-commerce platforms using a multimodal large model, abnormal content can be automatically located and edited, solving the problem of low efficiency of manual intervention in existing technologies and achieving efficient content review.
Patent Information
- Application Number
- CN202510961754.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies cannot accurately locate abnormal content in the multimodal description data of product information pages on e-commerce platforms, resulting in low efficiency of manual intervention and failing to meet the review requirements of e-commerce platforms.
Multimodal large model is used to process multimodal description data, determine the existence and type of abnormal content, and automatically edit abnormal content through location information and target processing methods to generate target description data that meets the review standards.
Abnormal content in the description information of target products can be accurately processed without human intervention, improving the review efficiency of e-commerce platforms and reducing costs.
Smart Images

Figure CN120876016A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to the fields of data processing, big data models and intelligent e-commerce technology, specifically to a content moderation method, apparatus, electronic device, computer-readable storage medium and computer program product. Background Technology
[0002] E-commerce platforms typically include multimodal descriptions of products on their product information pages, such as text, images, and videos. Before listing products, these multimodal descriptions must be reviewed to ensure the accuracy and compliance of the product information displayed on the product information page.
[0003] Currently, multimodal large models can be used to process the multimodal description data of related products in order to obtain content review results for the aforementioned multimodal description data.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a document generation method, apparatus, electronic device, computer-readable storage medium, and computer program product.
[0006] According to one aspect of this disclosure, a content review method is provided, comprising: acquiring multimodal description data associated with a target product, wherein the multimodal description data is associated with the description content of the target product on a product information page of an e-commerce platform; processing the multimodal description data using a multimodal big model to determine whether the multimodal description data includes abnormal content that does not meet review standards and the abnormal type of the abnormal content, wherein the abnormal type includes at least one of content inconsistency and the presence of illegal content; in response to determining that the multimodal description data includes the abnormal content, determining the location information of the abnormal content in the multimodal description data and a target processing method for the abnormal content based on the abnormal type, wherein the location information includes the position of the abnormal content on the product information page and at least one of the associated data modality; and processing the multimodal description data based on the target processing method to obtain target description data that meets the review standards.
[0007] According to another aspect of this disclosure, a first module is provided, configured to acquire multimodal description data associated with a target product, wherein the multimodal description data is associated with the description content of the target product on a product information page of an e-commerce platform; a second module is configured to process the multimodal description data using a multimodal big model to determine whether the multimodal description data includes abnormal content that does not meet the review standards and the abnormal type of the abnormal content, wherein the abnormal type includes at least one of content inconsistency and the presence of illegal content; a third module is configured to, in response to determining that the multimodal description data includes the abnormal content, determine the location information of the abnormal content in the multimodal description data and a target processing method for the abnormal content based on the abnormal type, wherein the location information includes at least one of the position of the abnormal content on the product information page and the associated data modality; and a fourth module is configured to process the multimodal description data based on the target processing method to obtain target description data that meets the review standards.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods described above.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the above-described method.
[0010] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the above-described method.
[0011] According to one or more embodiments of this disclosure, a content review method is provided. This method uses a multimodal model to review the multimodal description data associated with the product information page of a target product. It obtains a review result indicating whether abnormal content exists in the multimodal description data and the type of abnormality. When the review result indicates the presence of abnormal content, the method locates the position of the abnormal content on the product information page according to the type of abnormality and determines the corresponding target processing method. This allows for targeted editing of the abnormal content, ultimately resulting in target description data that meets the listing standards of the e-commerce platform. Therefore, abnormal content in the description information related to the target product can be accurately processed without manual intervention, effectively improving the review efficiency of the e-commerce platform and reducing costs.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 This is a schematic diagram illustrating an example system in which various methods described herein may be implemented according to exemplary embodiments;
[0015] Figure 2 A flowchart of a content moderation method according to an embodiment of the present disclosure is shown;
[0016] Figure 3 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0017] Figure 4 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0018] Figure 5 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0019] Figure 6 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0020] Figure 7 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0021] Figure 8 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0022] Figure 9 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0023] Figure 10 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0024] Figure 11 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0025] Figure 12 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0026] Figure 13 A partial flowchart of another content moderation method according to an embodiment of the present disclosure is shown;
[0027] Figure 14 A structural block diagram of a content moderation apparatus according to embodiments of the present disclosure is shown; and
[0028] Figure 15 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0030] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0031] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0032] E-commerce platforms typically include multimodal descriptions of products on their product information pages, such as text, images, and videos. Before listing products, these multimodal descriptions must be reviewed to ensure the accuracy and compliance of the product information displayed on the product information page.
[0033] In related technologies, multimodal large models can be used to process multimodal description data of related products in order to obtain content review results for the aforementioned multimodal description data.
[0034] However, multimodal large models typically determine review results by extracting semantic features and aligning feature consistency across multimodal data. Therefore, this approach cannot accurately pinpoint the specific location of anomalous content within the multimodal descriptive data (e.g., anomalous content may exist in image data, but its exact region within the image cannot be located). Consequently, currently, manual location and editing of rejected anomalous content is required, resulting in low efficiency.
[0035] To address the aforementioned issues, this disclosure provides a content moderation method. This method uses a multimodal big data model to review the multimodal description data associated with the product information page of a target product. It obtains a review result indicating whether abnormal content exists in the multimodal description data and the type of abnormality. When the review result indicates the presence of abnormal content, it locates the corresponding position on the product information page based on the type of abnormality and determines the appropriate target processing method. This allows for targeted editing of the abnormal content, ultimately resulting in target description data that meets the listing standards of e-commerce platforms. Therefore, abnormal content in the description information related to the target product can be accurately processed without manual intervention, effectively improving the review efficiency of e-commerce platforms and reducing costs.
[0036] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0037] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0038] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of content moderation methods.
[0039] In some embodiments, server 120 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.
[0040] exist Figure 1In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the content moderation methods described in this article, and is not intended to be restrictive.
[0041] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to perform content moderation methods. The client devices can provide interfaces that allow users to interact with them. The client devices can also output information to users through these interfaces. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0042] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows MobileOS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0043] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0044] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0045] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0046] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0047] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0048] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as text files, image files, and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0049] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0050] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0051] Figure 2 A flowchart of a content moderation method according to an embodiment of the present disclosure is shown.
[0052] like Figure 2 As shown, content moderation method 200 includes:
[0053] Step 210: Obtain multimodal description data associated with the target product, wherein the multimodal description data is associated with the description content of the target product on the product information page of the e-commerce platform;
[0054] Step 220: Use a multimodal large model to process the multimodal description data to determine whether the multimodal description data includes abnormal content that does not meet the review standards and the abnormal type of the abnormal content. The abnormal type includes at least one of content inconsistency and the presence of illegal content.
[0055] Step 230: In response to determining that the multimodal description data includes anomalous content, determine the location information of the anomalous content in the multimodal description data and the target processing method for the anomalous content based on the anomalous type, wherein the location information includes at least one of the anomalous content's position on the product information page and its associated data modality; and
[0056] Step 240: Process the multimodal description data based on the target processing method to obtain target description data that meets the review standards.
[0057] Therefore, by using a multimodal large model to review the multimodal description data associated with the product information page of the target product, the review results are obtained indicating whether there is abnormal content in the aforementioned multimodal description data and the type of abnormality. When the review results show that abnormal content exists, the location of the abnormal content in the product information page is located according to the type of abnormality, and the corresponding target processing method is determined. This allows for targeted editing of the abnormal content, ultimately resulting in target description data that meets the listing standards of the e-commerce platform. Abnormal content in the description information related to the target product can be accurately processed without manual intervention, effectively improving the review efficiency of the e-commerce platform and reducing costs.
[0058] In step 210, the multimodal description data may include at least two of text data, image data, and video data.
[0059] In step 210, the target product can be, for example, a product to be listed. If a merchant wants to list a new product on an e-commerce platform, they need to provide corresponding multimodal descriptive data based on the page style of the product information page on the e-commerce platform. For example, the merchant may need to fill in the title of the product to be listed, images, videos, and video data to display the details of the product, as well as product attribute information such as the material, size, and model of the product, so that the corresponding product information page can be generated and displayed on the e-commerce platform.
[0060] In step 220, the review criteria can be determined based on relevant laws and regulations and historical abnormal content data of the e-commerce platform.
[0061] In step 220, inconsistencies in content can be, for example, inconsistencies between the product title and the product details description, or inconsistencies between the text data and image data in the product details description; violations can be, for example, the presence of prohibited promotional language in the multimodal description data.
[0062] In step 240, e-commerce platforms typically need to review the multimodal description data submitted by merchants before allowing the corresponding products to be listed. However, to further improve efficiency, for some low-risk products (e.g., products with historical data showing few instances of piracy and false advertising), the products submitted by merchants can be listed first, and the multimodal description data can be obtained and the content reviewed only after a user views or places an order, thus saving computing resources.
[0063] In step 240, after obtaining the target description data, it can be returned to the merchant for confirmation. The data can then be listed after the merchant confirms that it is correct, in order to further improve the accuracy of the review and improve the processing effect.
[0064] Figure 3 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0065] According to some embodiments, such as Figure 3 As shown, step 220 includes:
[0066] Step 310: Extract semantic features from the multimodal description data to obtain multimodal semantic features;
[0067] Step 320: Map the multimodal semantic features to the target vector space for semantic consistency alignment to obtain the semantic alignment vector; and
[0068] Step 330: Calculate semantic alignment degree and identify illegal content based on semantic alignment vector to determine whether the multimodal description data includes abnormal content and abnormal type.
[0069] Therefore, by using a multimodal large model to process multimodal descriptive data, the characteristics of each data modality can be fully considered, as well as the semantic consistency between data of different modalities. This allows for efficient and accurate identification of whether there is abnormal content in the multimodal descriptive data, and the corresponding anomaly type when abnormal content exists.
[0070] In step 310, the multimodal semantic features are compared with the product attributes of the target product. For example, when the target product is clothing, the product attributes may include size, color, style, and material; when the target product is skincare product, the product attributes may include capacity, brand, usage method, and ingredients.
[0071] In step 320, since the data structures of different modalities are different, contrastive learning can be used to map the heterogeneous data of different modalities to the same high-dimensional vector space to achieve fine-grained alignment across modalities. For example, the high-dimensional vector space can be, for instance, the 512-dimensional vector space of the CLIP model.
[0072] In step 320, for example, semantic consistency alignment can be performed by aligning the textual semantic features of the text “pure cotton” with the second semantic features of the image showing the microfiber results of the fabric.
[0073] In step 320, dynamic associations between multimodal data can be established based on the cross-attention mechanism in the multimodal large model to generate semantic alignment vectors. For example, text data modality can be used as the primary modality, and image data and video data can be used as auxiliary modalities to better establish the aforementioned dynamic associations.
[0074] In step 330, the cosine similarity matrix between each multimodal feature can be calculated using semantic alignment vectors. If the similarity is found to be lower than a preset threshold (e.g., 0.4), it is determined that the two corresponding multimodal features are inconsistent and there is an anomaly.
[0075] In step 330, a keyword database of prohibited words can be pre-established for keyword detection, thereby identifying prohibited content and improving identification efficiency. For example, prohibited keywords could be phrases like "industry ceiling" and names of pirated brands, or similar prohibited promotional language.
[0076] For example, in scenarios involving counterfeit goods, the "open set recognition" + "text matching" model can be used to determine the review result, so as to effectively reduce the computational complexity in the processing.
[0077] In this example, the multimodal description data includes text and image data. A lightweight multimodal large model is used to extract visual open-set features from the image data, thereby outputting a set of candidate brand names (non-closed categories) for the target product as the first semantic feature of the image data. Then, a text large model is used to perform text similarity matching between the first semantic feature and the brand text in the brand database of the corresponding category to determine whether the current product is a counterfeit.
[0078] Specifically, when a merchant uploads a picture of sneakers with the letter "Aa" on the shoe, a multimodal large model can be used to identify and output "{"Aa","sneakers"} (open set recognition). Then, a text matching model is used to perform text similarity matching between "Aa" and the genuine "AA" brand in the brand database. If the similarity of the match is less than a preset threshold, the product is determined to be a counterfeit product and will not pass the review.
[0079] Based on the above method, the open set task only needs to identify the presence of brand identifiers, without needing to calculate high-dimensional classifications, effectively reducing computational complexity and improving overall response speed. Furthermore, using a large model for semantic recognition can more accurately identify the brand name of the product based on the overall context, improving the accuracy of the review results.
[0080] Figure 4 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0081] According to some embodiments, such as Figure 4 As shown, step 230 includes:
[0082] Step 410: In response to determining that the anomaly type includes content inconsistency, determine the first position of the first descriptive data associated with the anomaly content on the product information page and the second position of the second descriptive data on the product information page, wherein different positions on the product information page have different priority orders; and
[0083] Step 420: In response to determining that the priority of the first position is higher than the priority of the second position, determine the target processing method, which includes editing the second description data based on the first description data.
[0084] On a product information page, the density of information varies in different locations, which may also affect its credibility (for example, the title typically has lower information density, while the product details description usually has higher information density, and high-density information is generally more credible than low-density information). Therefore, for scenarios where the anomaly type is content inconsistency, different priorities can be set for different locations on the product information page. This allows for editing of low-priority anomaly content based on high-priority content, resulting in more accurate and compliant target description data.
[0085] In step 410, it is understood that the first and second description data here are only for illustrative purposes, and there may be more than two inconsistencies in the multimodal description data, which is not a limitation.
[0086] In step 410, the priority order of different positions on the product information page can be determined based on factors such as the information density of the corresponding position and the user's browsing order. For example, the title is usually the first position that users browse, so it is more likely to contain abnormal content, and its priority order is determined to be lower; while the product details description is usually the last position that users browse, so it is less likely to contain abnormal content, and its priority order is determined to be higher.
[0087] In step 420, for example, a large model can be used to process the first description data to obtain a semantic understanding result for the first description data, and then the second description data can be edited based on the semantic understanding result so that the edited second description data, while representing accurate information, meets the specific requirements of its location (e.g., word count limit requirements, etc.).
[0088] Figure 5 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0089] According to some embodiments, such as Figure 5 As shown, step 420 includes:
[0090] Step 510: In response to determining that the first position is located in the product details description section, and in response to determining that the second position is located in the product attribute description section or the product title section, determine that the priority of the first position is higher than the priority of the second position; and
[0091] Step 520: In response to determining that the priority of the first position is higher than the priority of the second position, the target processing method is determined to include:
[0092] Step 521: Extract multiple attribute features associated with the target product from the first description data, and generate a first knowledge graph based on the relationships between the multiple attribute features; and
[0093] Step 522: Edit the second description data based on the first knowledge graph.
[0094] The product details description typically includes highly credible and detailed product information, making it the highest priority reference section for editing data in other sections. Furthermore, due to the high information density of the product details description, a knowledge graph of the target product's attributes can be built based on it, allowing for simpler and more accurate editing of anomalous content in other sections.
[0095] In step 510, the product details description is more detailed, usually including all relevant information about the target product, such as product attributes, sales volume, awards, and other promotional language; while the product attribute description is less detailed, usually describing the various attributes of the product in a table format.
[0096] In step 520, for example, for a certain TV, the "screen" attribute in the product attributes section was mistakenly selected as "LCD (Liquid Crystal Display) screen", while the product details description section contained promotional terms such as "self-emissive pixels" and "infinite contrast ratio" that are actually related to OLED (Organic Light-Emitting Diode) screens.
[0097] At this point, a neural network model can be used to process the associated data in the product details description section, extract multiple attribute features associated with the TV, and construct a first knowledge graph based on the "category-attribute-promotional selling point" relationship between the extracted multiple attribute features. This allows matching the closest attribute node from the first knowledge graph based on the product attribute "screen". Since the multiple promotional selling point nodes associated with this attribute node describe OLED screens, the attribute node currently indicates "OLED screen". Therefore, it can be determined that the target processing method includes modifying the "screen" attribute in the product attribute section to "OLED screen".
[0098] According to one or more embodiments, the product attribute description section may have a higher priority than the product title. For example, if the product title uses the phrase "100% cotton," while the "composition" attribute in the product attribute description section describes "80% polyester fiber, 20% cotton," then the product title can be edited based on the product attribute description section to remove the phrase "100% cotton" or replace it with "not 100% cotton," etc.
[0099] Figure 6 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0100] According to some embodiments, such as Figure 6 As shown, the target processing methods also include:
[0101] Step 610: Obtain a second knowledge graph, wherein the second knowledge graph associates similar products of the target product, and the second knowledge graph indicates the association relationships between multiple second attribute features of similar products; and
[0102] Step 620: Edit the second description data based on the first knowledge graph and the second knowledge graph.
[0103] In some cases, relying solely on the product details description of the target product may not be enough to accurately determine the specific reference content used during editing. In such cases, a second knowledge graph of similar products in the same category as the target product can be introduced to supplement information and improve the accuracy of content editing.
[0104] For example, for the television in the above example, a second knowledge graph of other televisions with high attribute feature similarity can be obtained to supplement the content.
[0105] Figure 7 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0106] According to some embodiments, such as Figure 7 As shown, step 230 includes:
[0107] Step 710: In response to determining that the anomaly type includes content inconsistency, determine the first data modality of the third descriptive data and the second data modality of the fourth descriptive data associated with the anomaly content, wherein the data of different modalities in the product information page have different priority orders; and
[0108] Step 720: In response to determining that the priority of the first data modality is higher than that of the second data modality, determine the target processing method, which includes editing the fourth descriptive data based on the third descriptive data.
[0109] On product information pages, the credibility of content in different modalities of data may vary (for example, anomaly detection and identification in image and video data are more difficult, and therefore their credibility is lower than that of text data). Therefore, for scenarios where the anomaly type is content inconsistency, different priorities can be set for different modalities of data on the product information page. This allows for editing of anomaly content in lower-priority areas based on higher-priority anomalies, resulting in more accurate and compliant target description data.
[0110] In step 710, it is understood that the third and fourth description data here are only for illustrative purposes, and there may be more than two inconsistencies in the multimodal description data, which is not a limitation.
[0111] In step 710, the priority ranking of different positions on the product information page can be determined based on factors such as the computational complexity of detecting the corresponding data modality and the degree of content visualization. For example, compared to text, users may prefer to understand products through intuitive methods such as images and videos, so the possibility of their text data containing abnormal content is lower, and its priority is determined to be higher.
[0112] In step 720, for example, a large model can be used to process the third descriptive data to obtain a semantic understanding result for the first descriptive data, and then the second descriptive data can be edited based on the semantic understanding result so that the edited second descriptive data, while representing accurate information, meets the specific requirements of its location (e.g., data modality and word count limits).
[0113] Figure 8 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0114] According to some embodiments, such as Figure 8 As shown, step 720 includes:
[0115] Step 810: In response to determining that the first data modality is a text data modality, and in response to determining that the second data modality is an image data modality, determine that the priority of the first data modality is higher than the priority of the second data modality; and
[0116] Step 820: In response to determining that the priority of the first data mode is higher than the priority of the second data mode, the target processing method is determined to include:
[0117] Step 821: Determine the product attributes of the target product associated with the third descriptive data, and calculate the attention heatmap of the fourth descriptive data based on the product attributes to obtain the processing area of the associated product attributes; and
[0118] Step 822: Edit the area to be processed based on the third description data.
[0119] Text data typically contains highly reliable and detailed product information, and its processing is computationally inefficient and highly accurate. Therefore, it can be used as the highest priority reference data, upon which other data modalities can be edited. Furthermore, by calculating attention heatmaps of associated product data, the regions containing anomalous content in image data can be identified more effectively and quickly, thus improving processing efficiency.
[0120] In one example, the area to be processed includes text content that is inconsistent with the third description data. In this case, the text content in the area to be processed can be replaced or eliminated using the content from the third description data. The method for processing the text content can be referred to the explanation for step 420 above, and will not be repeated here.
[0121] In one example, the area to be processed includes non-text content, such as "green clothes". However, according to the third description data, the target product should be "red clothes". Therefore, object recognition and image processing can be performed based on the third description data to modify the color of the clothes in the area to be processed.
[0122] According to one or more embodiments, to further expand the scope of application, the target processing method can also be determined simultaneously based on the position of the abnormal content on the product information page and the associated data modality. For example, the fourth and fifth positions of the abnormal content on the product information page can be determined first. In response to determining that the third and fourth positions have the same priority, the third data modality associated with the fourth position and the fourth data modality associated with the fifth position can be determined. Then, the target processing method can be determined based on the priority relationship between the third and fourth data modalities.
[0123] Figure 9 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0124] According to some embodiments, such as Figure 9 As shown, the multimodal description data includes first video data, and the third description data is the first anomalous image frame in the first video data. Step 240 includes:
[0125] Step 910: Obtain the first abnormal image frame from the first video data; and
[0126] Step 920: Process the first abnormal image frame based on the target processing method to update the first video data, thereby obtaining the target description data.
[0127] For cases where multimodal description data includes video data, the video data can be divided into multiple image frames, and abnormal image frames can be identified and processed as described above to improve processing efficiency and expand the applicability of the above method.
[0128] Figure 10 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0129] According to some embodiments, such as Figure 10 As shown, step 230 includes:
[0130] Step 1010: In response to determining the anomaly type, including the presence of illegal content, determine the third location of the anomaly content based on the location information; and
[0131] Step 1020: Determine the target processing method, including deleting the violating content from the third location.
[0132] Therefore, for cases where the anomaly type indicates the presence of illegal content, the corresponding illegal content can be directly deleted from the multimodal description data to improve review efficiency and reduce costs.
[0133] In step 1020, if it is determined that the data modality associated with the third position is text data, the illegal text content can be directly deleted from the third position to improve processing efficiency.
[0134] For example, OCR (Optical Character Recognition) can be used to locate the coordinates of abnormal content, and then edit it.
[0135] Figure 11 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0136] According to some embodiments, such as Figure 11 As shown, step 1020 includes:
[0137] Step 1110: Process image data using an image recognition model to identify the target region where the target object associated with the abnormal content is located from the image data; and
[0138] Step 1120: Use a grayscale mask to cover the target area to eliminate abnormal content.
[0139] Therefore, by using grayscale masks, abnormal content in image data can be covered simply and efficiently, thereby improving processing efficiency and reducing processing difficulty.
[0140] For example, abnormal content can be a forged certificate, and the target area can be the area where the forged certificate is located; for example, abnormal content can be a spokesperson. For instance, if the audit results determine that the target product is not associated with any spokesperson, but the image data contains a person's image and the person's image is marked with the words "Spokesperson: AAA", then the area can be determined as the target area, which has a high risk of violation and needs to be covered.
[0141] Figure 12 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0142] According to some embodiments, such as Figure 12 As shown, after step 1120, method 200 further includes:
[0143] Step 1210: Obtain the image data after removing abnormal content as the image data to be repaired; and
[0144] Step 1220: Process the image data to be repaired using an image edge repair model, wherein the image edge repair model is used to improve the image fusion between the grayscale mask and the image data at the edges of the target area.
[0145] Therefore, after removing abnormal content from image data, the edges of the removed area can be further processed to enhance image fusion, thereby improving the overall display effect of the image data and further enhancing the user experience.
[0146] In step 1220, for example, EdgeConnect Generative Adversarial Network (a partially convolutional generative adversarial network structure) can be used for edge inpainting. Specifically, the U-Net structure (a network structure) in the network model is used as a generator to inpaint image edges, and PatchGAN (a network structure) is used as a discriminator to determine whether the fusion degree of the inpainted region meets the requirements. Furthermore, a loss function can be established based on L1 content loss, style loss, and adversarial loss to enable model convergence.
[0147] Based on the above model, the limitations of traditional mosaic and blur processing can be overcome, and the background texture can be adaptively reconstructed, edge continuity can be maintained (edge matching degree greater than 85%), and visual perception can be seamless (structural similarity index greater than 0.92) can be achieved during the edge repair process.
[0148] Figure 13 A partial flowchart of another content moderation method according to an embodiment of this disclosure is shown.
[0149] According to some embodiments, such as Figure 13As shown, the multimodal description data includes the second video data, and the third description data is the first anomalous image frame in the first video data. Step 240 includes:
[0150] Step 1310: Obtain the second abnormal image frame from the second video data; and
[0151] Step 1320: Process the second abnormal image frame based on the target processing method to update the second video data, thereby obtaining the target description data.
[0152] For cases where multimodal description data includes video data, the video data can be divided into multiple image frames, and abnormal image frames can be identified and processed as described above to improve processing efficiency and expand the applicability of the above method.
[0153] According to another aspect of this disclosure, a content moderation device is provided. For example... Figure 14 As shown, the content review device 1400 includes: a first module 1410 configured to acquire multimodal description data associated with a target product, wherein the multimodal description data is associated with the description content of the target product on the product information page of an e-commerce platform; a second module 1420 configured to process the multimodal description data using a multimodal big model to determine whether the multimodal description data includes abnormal content that does not meet the review standards and the abnormal type of the abnormal content, wherein the abnormal type includes at least one of content inconsistency and the presence of illegal content; a third module 1430 configured to, in response to determining that the multimodal description data includes abnormal content, determine the location information of the abnormal content in the multimodal description data and the target processing method for the abnormal content based on the abnormal type, wherein the location information includes the position of the abnormal content on the product information page and at least one of the associated data modality; and a fourth module 1440 configured to process the multimodal description data based on the target processing method to obtain target description data that meets the review standards.
[0154] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the aforementioned method.
[0155] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause the computer to perform the aforementioned method.
[0156] According to another aspect of this disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements the aforementioned method when executed by a processor.
[0157] like Figure 15 As shown, the electronic device 1500 includes a computing unit 1501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1502 or a computer program loaded from a storage unit 1508 into a random access memory (RAM) 1503. The RAM 1503 may also store various programs and data required for the operation of the electronic device 1500. The computing unit 1501, ROM 1502, and RAM 1503 are interconnected via a bus 1504. An input / output (I / O) interface 1505 is also connected to the bus 1504.
[0158] Multiple components in electronic device 1500 are connected to I / O interface 1505, including: input unit 1506, output unit 1507, storage unit 1508, and communication unit 1509. Input unit 1506 can be any type of device capable of inputting information to electronic device 1500. Input unit 1506 can receive input digital or character information and generate key signal input related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 1507 can be any type of device capable of presenting information, and can include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1508 can include, but is not limited to, disk and optical disk. Communication unit 1509 allows electronic device 1500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets such as Bluetooth. TM Devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0159] The computing unit 1501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1501 performs the various methods and processes described above, such as content moderation methods. For example, in some embodiments, the content moderation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1500 via ROM 1502 and / or communication unit 1509. When the computer program is loaded into RAM 1503 and executed by the computing unit 1501, one or more steps of the content moderation method described above may be performed. Alternatively, in other embodiments, the computing unit 1501 may be configured to perform GPU-based matrix computation methods by any other suitable means (e.g., by means of firmware).
[0160] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0161] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0163] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0164] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0165] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0166] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0167] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A content moderation method, comprising: Obtain multimodal description data associated with a target product, wherein the multimodal description data is associated with the description content of the target product on the product information page of an e-commerce platform; The multimodal large model is used to process the multimodal description data to determine whether the multimodal description data includes abnormal content that does not meet the review standards and the abnormal type of the abnormal content, wherein the abnormal type includes at least one of content inconsistency and the presence of illegal content; In response to determining that the multimodal description data includes the anomalous content, the location information of the anomalous content in the multimodal description data and the target processing method for the anomalous content are determined according to the anomalous type, wherein the location information includes at least one of the position of the anomalous content on the product information page and the associated data modality; and The multimodal description data is processed based on the target processing method to obtain target description data that meets the review criteria.
2. The method according to claim 1, wherein, The step of determining the location information of the abnormal content in the multimodal description data and the target processing method for the abnormal content based on the abnormality type includes: In response to determining that the anomaly type includes the content inconsistency, the system determines a first position of first descriptive data associated with the anomaly content on the product information page and a second position of second descriptive data associated with the anomaly content on the product information page, wherein the different positions on the product information page have different priority orders; and In response to determining that the priority of the first position is higher than the priority of the second position, the target processing method is determined to include editing the second description data based on the first description data.
3. The method according to claim 2, wherein, The step of determining the target processing method in response to determining that the priority of the first position is higher than the priority of the second position includes editing the second description data based on the first description data, including: In response to determining that the first position is located in the product details description section, and in response to determining that the second position is located in the product attribute description section or the product title section, the priority of the first position is determined to be higher than the priority of the second position; and In response to determining that the priority of the first position is higher than the priority of the second position, determining the target processing method includes: Extract multiple attribute features associated with the target product from the first description data, and generate a first knowledge graph based on the relationships between the multiple attribute features; and The second description data is edited based on the first knowledge graph.
4. The method according to claim 3, wherein the target processing method further includes: Obtain a second knowledge graph, wherein the second knowledge graph associates similar products of the target product, and the second knowledge graph indicates the association relationships between multiple second attribute features of the similar products; and The second description data is edited based on the first knowledge graph and the second knowledge graph.
5. The method according to claim 1, wherein, The step of determining the location information of the abnormal content in the multimodal description data and the target processing method for the abnormal content based on the abnormality type includes: In response to determining that the anomaly type includes the content inconsistency, a first data modality of third descriptive data and a second data modality of fourth descriptive data associated with the anomaly content are determined, wherein the data of different modalities in the product information page have different priority orders; and In response to determining that the priority of the first data modality is higher than that of the second data modality, the target processing method is determined to include editing the fourth description data based on the third description data.
6. The method according to claim 5, wherein, In response to determining that the priority of the first data modality is higher than the priority of the second data modality, determining the target processing method includes editing the fourth description data based on the third description data, including: In response to determining that the first data modality is a text data modality, and in response to determining that the second data modality is an image data modality, it is determined that the priority of the first data modality is higher than the priority of the second data modality; and In response to determining that the priority of the first data modality is higher than the priority of the second data modality, determining the target processing method includes: The product attributes of the target product associated with the third descriptive data are determined, and an attention heatmap of the fourth descriptive data is calculated based on the product attributes to obtain the processing region associated with the product attributes; and The region to be processed is edited based on the third descriptive data.
7. The method according to claim 6, wherein, The multimodal description data includes first video data, and the third description data is a first abnormal image frame in the first video data. Processing the multimodal description data based on the target processing method to obtain target description data that conforms to the review criteria includes: Obtain the first abnormal image frame from the first video data; and The first abnormal image frame is processed based on the target processing method to update the first video data, thereby obtaining the target description data.
8. The method according to claim 1, wherein, The step of determining the location information of the abnormal content in the multimodal description data and the target processing method for the abnormal content based on the abnormality type includes: In response to determining that the anomaly type includes the presence of illegal content, a third location of the anomaly content is determined based on the location information; and The target processing method is determined to include deleting the violating content from the third location.
9. The method according to claim 8, wherein, The third location is associated with image data, and the step of removing the abnormal content from the third location includes: The image data is processed using an image recognition model to identify the target region where the target object associated with the abnormal content is located; and Use a grayscale mask to cover the target area to eliminate the abnormal content.
10. The method according to claim 9, wherein, After covering the target area with a grayscale mask, the method further includes: Obtain the image data after removing the abnormal content as the image data to be repaired; and The image data to be repaired is processed using an image edge restoration model, wherein the image edge restoration model is used to improve the image blending degree between the grayscale mask and the image data at the edge of the target region.
11. The method according to any one of claims 8-10, wherein, The multimodal description data includes second video data, and the image data is a second abnormal image frame in the second video data. The step of processing the multimodal description data based on the target processing method to obtain target description data that conforms to the review criteria includes: Obtain the second abnormal image frame from the second video data; and The second abnormal image frame is processed based on the target processing method to update the second video data, thereby obtaining the target description data.
12. The method according to any one of claims 1-11, wherein, The process of using a multimodal large model to process the multimodal description data includes: Semantic features are extracted from the multimodal description data to obtain multimodal semantic features; The multimodal semantic features are mapped to the target vector space for semantic consistency alignment, resulting in a semantic alignment vector; and Based on the semantic alignment vector, semantic alignment degree calculation and violation content identification are performed to determine whether the multimodal description data includes the abnormal content and the abnormal type.
13. A content moderation device, comprising: The first module is configured to acquire multimodal description data associated with the target product, wherein the multimodal description data is associated with the description content of the target product in the product information page of the e-commerce platform; The second module is configured to process the multimodal description data using a multimodal large model to determine whether the multimodal description data includes abnormal content that does not meet the review standards and the abnormal type of the abnormal content, wherein the abnormal type includes at least one of content inconsistency and the presence of illegal content; The third module is configured to, in response to determining that the multimodal description data includes the anomalous content, determine, based on the anomalous type, the location information of the anomalous content in the multimodal description data and a target processing method for the anomalous content, wherein the location information includes at least one of the location of the anomalous content on the product information page and the associated data modality; and The fourth module is configured to process the multimodal description data based on the target processing method to obtain target description data that meets the review criteria.
14. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-12.
16. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-12.