Infringement commodity detection method and device, equipment and medium
By introducing pre-infringement detection into the product listing process of e-commerce platforms, combined with text and image analysis, the problem of frequent infringement of product details pages on e-commerce platforms is solved, efficient and accurate infringement identification and timely blocking is achieved, and trademarks and copyrights are protected.
Patent Information
- Application Number
- CN202510948720.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-10
AI Technical Summary
In the prior art, infringement of product details pages on e-commerce platforms frequently occurs, traditional detection methods are inefficient and insufficiently accurate, making it difficult to effectively identify infringement under similar semantic or complex backgrounds.
By introducing a pre-infringement detection mechanism in the product listing process, combining natural language processing and image recognition technology, analyzing product description text and pictures, judging trademark and copyright infringement, generating infringement detection reports and preventing listing.
It has realized comprehensive infringement detection on product details pages, improved the accuracy and comprehensiveness of the detection, timely discovered and prevented potential infringements, protected intellectual property rights and maintained platform compliance.
Smart Images

Figure CN120450828A_ABST
Abstract
Description
Technical Field
[0001] This application covers the field of e-commerce technology, and in particular includes a method for detecting infringing goods and its device, equipment, and medium. Background Art
[0002] Against the backdrop of the rapid development of e-commerce, infringements on product details pages have gradually become an urgent issue to be addressed. With the popularization of e-commerce platforms, the display and dissemination of product information has become more convenient, but this has also led to frequent trademark and copyright infringements. These infringements not only damage the legitimate rights and interests of brands and creators, but also undermine the fair competition environment in the market.
[0003] Traditional methods for detecting infringing goods rely primarily on manual review or single-dimensional technical means. While manual review can make relatively accurate judgments about infringement, its low efficiency makes it difficult to meet the needs of rapidly reviewing large amounts of product information. Furthermore, manual review is susceptible to subjective factors, leading to inconsistent review results. On the other hand, single-dimensional technical means, such as simple string matching or rule-based image recognition, can improve the efficiency of infringement detection to a certain extent, but their accuracy and comprehensiveness still need to be improved. For example, simple string matching methods can only detect identical text content and are powerless against infringements with similar semantics but different expressions. Rule-based image recognition methods, on the other hand, struggle to cope with complex image backgrounds and diverse trademark styles, and are prone to misjudgments or omissions. Summary of the Invention
[0004] The primary purpose of this application is to solve at least one of the above problems and provide a method for detecting infringing goods and its device, equipment, and medium.
[0005] According to one aspect of the present application, a method for detecting infringing goods is provided, comprising the following steps: In response to a user-triggered target product listing event, the pre-infringement detection process in the listing process of the target product is initiated, and the product description text and product image on the product details page of the target product are obtained; Determine whether the product description text contains a word whose similarity to the preset trademark text information is greater than a preset similarity threshold, or whether the product image contains a preset trademark brand logo. If so, determine that the product details page constitutes trademark infringement; Determining the relevance between the product description text and product image and a preset set of copyright description text and copyright image. If any of the relevances meets a preset condition, the product details page is determined to constitute copyright infringement. When the product details page is judged to be a trademark infringement or copyright infringement, a corresponding infringement detection report is constructed and pushed to the user, ending the listing process of the target product.
[0006] According to another aspect of the present application, a device for detecting infringing goods is provided, comprising: An event response module is configured to respond to a user-triggered target product listing event, initiate a pre-infringement detection process in the target product listing process, and obtain product description text and product images from the target product's product details page; a trademark infringement detection module configured to determine whether the product description text contains a word whose similarity to a preset trademark text information is greater than a preset similarity threshold, or whether the product image contains a preset trademark brand logo, and if so, determine that the product details page constitutes trademark infringement; a copyright infringement detection module configured to determine the relevance between the product description text and product image and a preset set of copyright description text and copyright image, and to determine that the product details page constitutes copyright infringement when any of the relevances meets a preset condition; The result processing module is configured to construct a corresponding infringement detection report and push it to the user when the product details page is judged to be a trademark infringement or copyright infringement, thereby ending the listing process of the target product.
[0007] According to another aspect of the present application, an infringing goods detection device is provided, comprising a central processing unit and a memory, wherein the central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the infringing goods detection method described in the present application.
[0008] According to another aspect of the present application, a non-volatile readable storage medium is provided, which stores a computer program implemented according to the infringing goods detection method in the form of computer-readable instructions. When the computer program is called and executed by a computer, the steps included in the method are executed.
[0009] The technical solution of this application has many advantages, including but not limited to the following: This application, by pre-positioning an infringement detection mechanism in the product listing process, can promptly detect and prevent potential infringements before the products are officially listed, thus avoiding the discovery of infringing products only after they have been widely disseminated on the platform, thereby minimizing damage to trademark and copyright owners, while also maintaining the business reputation and legal compliance of the e-commerce platform.
[0010] Specifically, this application determines whether there is trademark infringement by judging whether there are word segments in the product description text whose similarity with the preset trademark text information exceeds a threshold, and whether the product image contains the preset trademark brand logo. At the same time, by confirming the correlation between the product description text and product image and the preset copyright description text set and copyright image set, it determines whether there is copyright infringement. When the product details page is judged to be a trademark infringement or copyright infringement, an infringement detection report is constructed and pushed to the user, and the product listing process is terminated at the same time. That is, this application realizes comprehensive infringement detection of product details pages, which not only covers the product description text, but also includes product images. By comprehensively analyzing the infringement information in the text and images, it can more comprehensively identify trademark infringement and copyright infringement. This multi-dimensional detection method significantly improves the accuracy and comprehensiveness of infringement detection, and avoids infringements that may be missed by a single-dimensional detection method. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 The network architecture of the e-commerce platform exemplified in this application; Figure 2 A flowchart of a typical embodiment of the infringing goods detection method of the present application; Figure 3 This is a principle block diagram of the infringing goods detection device of this application; Figure 4 This is a schematic diagram of the structure of an infringing product detection device used in this application. DETAILED DESCRIPTION
[0012] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.
[0013] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0014] Those skilled in the art will understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless specifically defined as such, will not be interpreted in an idealized or overly formal sense.
[0015] like Figure 1 In the network architecture shown, the e-commerce platform 82 is deployed on the Internet to provide corresponding services to its users. Similarly, the devices 80 of the merchant users and the devices 81 of the consumer users of the e-commerce platform 82 are also connected to the Internet to use the services provided by the e-commerce platform.
[0016] The exemplary e-commerce platform 82 provides supply and demand matching of products and / or services to the general public with the help of Internet infrastructure. In the e-commerce platform 82, products and / or services are provided as commodity information. To simplify the description, the concepts of commodity, product, etc. are used in this application to refer to the products and / or services in the e-commerce platform 82, which may specifically be physical products, digital products, tickets, service subscriptions, other offline services, etc.
[0017] In reality, various entities can access the e-commerce platform 82 as users, use the various online services provided by the e-commerce platform 82, and achieve the purpose of participating in the business activities achieved by the e-commerce platform 82. These entities can be natural persons, legal persons, or social organizations. Corresponding to the two types of entities in business activities, merchants and consumers, the e-commerce platform 82 has two corresponding types of users: merchant users and consumer users. In business activities, all entities in the product distribution chain, including manufacturers, sellers, retailers, logistics providers, etc., can use online services on the e-commerce platform 82 as merchant users, while consumers in business activities, including real or potential consumers, can use online services on the e-commerce platform 82 as their corresponding consumer users. In actual business activities, the same entity can act as both a merchant user and a consumer user, and this should be understood flexibly.
[0018] The infrastructure used to deploy the e-commerce platform 82 primarily includes a backend architecture and frontend devices. The backend architecture runs various online services through a service cluster, including platform-facing middleware or frontend services, consumer-facing services, merchant-facing services, etc., to enrich and improve its service functions. The frontend devices primarily encompass the terminal devices used by users to access the e-commerce platform 82 as clients, including but not limited to various mobile terminals, personal computers, point-of-sale devices, etc. For example, a merchant user can use their terminal device 80 to enter product information for their online store, or generate their product information using an interface open to the e-commerce platform. A consumer user can use their terminal device 81 to access the webpage of the online store implemented by the e-commerce platform 82, trigger the shopping process through the shopping button provided on the webpage, and invoke various online services provided by the e-commerce platform 82 during the shopping process, thereby completing the purpose of placing a shopping order.
[0019] In some embodiments, the e-commerce platform 82 may be implemented by a processing facility including a processor and a memory, the processing facility storing a set of instructions that, when executed, cause the e-commerce platform 82 to perform the e-commerce and support functions described herein. The processing facility may be part of a server, client, network infrastructure, mobile computing platform, cloud computing platform, fixed computing platform, or other computing platform, and may provide electronic components of the e-commerce platform 82, merchant devices, payment gateways, application developers, marketing channels, transportation providers, customer devices, point-of-sale devices, and the like.
[0020] The e-commerce platform 82 can be implemented as an online service such as cloud computing, software as a service (SaaS), infrastructure as a service (IaaS), platform as a service (PaaS), desktop as a service (DaaS), hosted software as a service, mobile backend as a service (MBaaS), and information technology management as a service (ITMaaS). In some embodiments, the various functional components of the e-commerce platform 82 can be implemented to operate on various platforms and operating systems. For example, for an online store, its administrator users can enjoy the same or similar functions regardless of the various embodiments, such as iOS, Android, HomonyOS, or the web.
[0021] The e-commerce platform 82 can implement its corresponding independent website for each merchant to run its corresponding online store, and provide merchants with corresponding business management engine instances for merchants to establish, maintain, and run one or more online stores in one or more independent websites. The business management engine instance can be used for content management, task automation, and data management of one or more online stores, and can configure various specific business processes of the online store through interfaces or built-in components to support the implementation of business activities. The independent website is the infrastructure of the e-commerce platform 82 with cross-border service functions. Merchants can maintain their online stores more centrally and independently based on the independent website. The independent website usually has a domain name and storage space dedicated to the merchant, and different independent websites are relatively independent. The e-commerce platform 82 can provide standardized or personalized technical support for a large number of independent websites, so that merchant users can customize their own business management engine instance and use this business management engine instance to maintain one or more online stores they own.
[0022] The online store can implement backend configuration and maintenance by having the merchant user log in to its business management engine instance as an administrator. With the support of various online services provided by the infrastructure of the e-commerce platform 82, the merchant user can configure various functions in its online store as an administrator, view various data, etc. For example, the merchant user can manage various aspects of its online store, such as viewing the latest activities of the online store, updating the online store product catalog, managing orders, recent visit activities, total order activities, etc.; the merchant user can also view more detailed information about the business and visitors to the merchant's online store by obtaining reports or metrics, such as showing a sales summary of the merchant's overall business, specific sales and participation data of active sales marketing channels, etc.
[0023] The e-commerce platform 82 may provide communication facilities and associated merchant interfaces for providing electronic communications and marketing, such as utilizing electronic message aggregation facilities to collect and analyze communication interactions between merchants, consumers, merchant devices, customer devices, point-of-sale devices, etc., and to aggregate and analyze communications, such as for increasing the potential for providing product sales, etc. For example, a consumer may have questions about a product, which may generate a conversation between the consumer and the merchant (or an automated processor-based agent on behalf of the merchant), wherein the communication facility is responsible for the interaction and provides the merchant with analysis on how to increase the probability of a sale.
[0024] In some embodiments, applications suitable for installation on terminal devices can be provided to serve the access needs of different users, so that various users can access the e-commerce platform 82 by running applications on the terminal devices, such as the merchant backend module of the online store in the e-commerce platform 82. In the process of implementing business activities through these functions, the e-commerce platform 82 can implement various functions related to supporting business activities as middleware or online services and open corresponding interfaces, and then implant toolkits corresponding to the interface access functions into the application to realize functional expansion and task implementation. The business management engine can include a series of basic functions and expose these functions to online services and / or application calls through APIs. The online services and applications use the corresponding functions by remotely calling the corresponding APIs.
[0025] Supported by the various components of the business management engine instance, the e-commerce platform 82 provides online shopping functionality, enabling merchants to connect with customers in a flexible and transparent manner. Consumers can select items online, create an order, provide a delivery address in the order, and complete payment confirmation for the order. Merchants can then review and complete or cancel the order. The review component included in the business management engine instance enables compliance with business processes to ensure that orders are suitable for fulfillment before actual fulfillment. Orders may sometimes be fraudulent and require verification. A payment method that requires merchants to wait to ensure receipt of funds can help mitigate this risk, and so on. Order risks may be generated by fraud detection tools submitted by third parties through, for example, the Order Risk API. Before fulfillment, merchants may need to obtain or wait for payment information to mark the order as paid, at which point the merchant can prepare to deliver the product. These situations can all be subject to appropriate review. The review process can be implemented by the fulfillment component. Merchants can use the fulfillment component to review and adjust their work and trigger relevant fulfillment services. For example, manual fulfillment is used when merchants pick and pack products in boxes, purchase shipping labels and enter their tracking numbers, or simply mark items as fulfilled; custom fulfillment can be used to define email notifications; API fulfillment can trigger third-party applications to create fulfillment records on third parties; legacy fulfillment can trigger custom API calls from the commerce management engine to third parties; and gift card fulfillment can generate numbers and activate gift cards. Merchants can use the order printer application to print shipping labels. Fulfillment can be executed when the items are packed in boxes and ready for shipping, tracking, delivery, and consumer verification of receipt.
[0026] It can be seen that the services provided by e-commerce platforms are precisely product-centric. The corresponding product data is the foundational data of the e-commerce platform. Providing product information through product data and mining and utilizing product data are the foundation for various technical services, including leveraging user transaction data and product data within the e-commerce platform's product data to provide fundamental services for the operation of the data processing system. Therefore, the data processing system can run on any one or multiple servers within the e-commerce platform's fleet, utilizing the various product data provided by the e-commerce platform to implement various functions.
[0027] See also Figure 2 The infringing goods detection method of the present application, in its typical embodiment, includes the following steps: Step S5100: In response to a user-triggered target product listing event, initiate a pre-infringement detection process in the target product listing process, and obtain the product description text and product image on the product details page of the target product; During the operation of the e-commerce platform, merchants need to upload product information to the product details page of the e-commerce platform so that consumers can browse and purchase. However, with the continuous increase in the number of products on the e-commerce platform, the number of infringements occurring in the product details page is also gradually increasing. In order to effectively prevent infringements in the product details page and ensure the healthy operation of the e-commerce platform, this application will start the pre-infringement detection process before the product is put on the shelf to perform infringement detection on the product details page corresponding to the product to be put on the shelf. That is, when the merchant triggers the listing event of the target product on the e-commerce platform, the infringing product detection method of this application is automatically triggered.
[0028] Specifically, the first step is to obtain the product description text and product images from the target product's product details page. The product description text is the merchant's detailed description of the product, including key information such as the product title, features, instructions for use, specifications, and brand information. The product description text is a key basis for determining trademark infringement. Since trademark infringement often involves the unauthorized use of another's trademark text or logo, natural language processing technology can be used to analyze the obtained product description text, extract key information, and compare it with pre-set trademark text information to determine whether there is a possibility of trademark infringement. Product images are important content that showcases the target product's appearance, including images of the physical product, packaging, and usage scenarios. Image infringement detection primarily targets the unauthorized use of copyrighted images or the unauthorized use of trademark logos in product images. In this case, image recognition technology can be used to analyze product images, extract key features, and compare them with a pre-set set of copyrighted images and trademark brand logos to determine whether there is infringement.
[0029] Through this pre-infringement detection mechanism, e-commerce platforms can discover potential infringements in the early stages of product listing and take timely measures, such as preventing products from being listed and notifying merchants to modify information, thereby effectively protecting intellectual property rights and maintaining the platform's business order.
[0030] Step S5200: Determine whether the product description text contains a word whose similarity to a preset trademark text information is greater than a preset similarity threshold, or whether the product image contains a preset trademark brand logo. If so, determine that the product details page constitutes trademark infringement. First, we conduct in-depth text analysis of product descriptions. This includes word segmentation, which breaks down sentences and paragraphs in product descriptions into individual lexical units. Trademark infringement often manifests itself in unauthorized use of keywords or phrases that represent another person's trademark. Word segmentation can help more accurately identify these potentially infringing terms. For example, if the target product description contains the brand name "Brand A," which is already a registered trademark, "Brand A" will be identified.
[0031] The segmented words are compared with the preset trademark text information for similarity. The product text information is stored in the product text information database. The trademark text information database is a database containing registered trademark keywords. The similarity between each word in the product description text and the words in the trademark text information database is calculated. The similarity can be calculated using a variety of algorithms, such as the cosine similarity algorithm, which measures the degree of similarity between the two by calculating the cosine value of the angle between two text vectors. If the similarity between a word and a trademark word in the trademark information database exceeds the preset similarity threshold, it will be marked as a potential infringing word. For example, if the preset similarity threshold is 0.8, and the similarity between the segmented word corresponding to the product description text and the "Brand A" trademark in the trademark information database is 0.9, it is determined that there is trademark infringement in the product description text.
[0032] In image recognition, trademark and brand logo detection is performed on product images. In one embodiment, a set of sample images containing trademark and brand logos is obtained. These sample images are carefully annotated, including information such as the location, shape, and features of the trademarks. Based on these annotated sample images, a logo detection model is trained to identify the presence of a preset trademark and brand logo in the product image. For example, if the "Brand A" checkmark logo appears in an image of the target product, the model will detect the presence of the checkmark logo. After the logo detection model detects the presence of the trademark and brand logo in the product image, it further determines the confidence level of the logo. The confidence level is a measure of the reliability of the model's detection results, a value between 0 and 1. If the confidence level is greater than or equal to a preset confidence threshold, such as 0.95, the product image is directly judged to have infringed the trademark. However, if the confidence level is lower than the preset threshold, further verification is required to determine whether the text information in the product image contains the trademark text information. Considering that the trademark logo in the product image may not be clear or there may be certain interference, resulting in low confidence in model detection, at this time, by combining verification with text information, it is possible to more accurately determine whether there is trademark infringement.
[0033] Through the above-mentioned comprehensive analysis of product description text and product images, it is possible to comprehensively determine whether the target product details page constitutes trademark infringement. This multi-dimensional detection method not only improves the accuracy of infringement detection, but also can effectively deal with various complex infringement behaviors.
[0034] Step S5300: confirming the correlation between the product description text and product image and a preset set of copyright description text and copyright image. If any correlation satisfies a preset condition, it is determined that the product details page constitutes copyright infringement. First, a copyright relevance analysis is performed on the product description text. The copyright description text set is a database containing copyrighted text content, including product titles, product details, or brand information from individual product detail pages. The target product's product description text is compared with the copyright description text set to determine whether there is unauthorized similar or identical content. Specifically, the product description text is compared with the corresponding copyright description text in the copyright description text set. During this comparison, natural language processing techniques are used to semantically vectorize the product and copyright description texts. Semantic vectorization is a technique that converts text content into numerical vectors, capturing the semantic features of text and thus enabling quantitative analysis of text similarity. The similarity between the product description text and the copyright description text is calculated. The preset condition is that if the similarity exceeds a preset text similarity threshold, for example, 0.7, a high correlation is determined between the product description text and the copyright description text, and the product detail page is deemed to constitute copyright infringement.
[0035] Furthermore, copyright relevance detection is performed on product images. The copyright image collection is a database containing image content that is protected by copyright. The copyright images in the database can be composed of various product images uploaded from various product detail pages. The product images of the target product are compared with the copyright image collection to determine whether there are unauthorized similar or identical images. During the image comparison process, image feature extraction technology is used to extract feature vectors of product images and copyright images. Image feature extraction technology can capture the visual features of images, such as color distribution, texture, shape, etc., and calculate the similarity between product images and copyright images. For example, by calculating the Euclidean distance or Hamming distance between the two image feature vectors, the preset condition is set to if the similarity exceeds a preset image similarity threshold, such as 0.8, it is determined that there is a high correlation between the product image and the copyright image, and it is determined that the product detail page constitutes copyright infringement.
[0036] Comprehensively consider the relevance of product descriptions and images. If the relevance between any of the product descriptions or images and the copyright description text set or copyright image set meets the preset criteria, the product details page is deemed to have infringed copyright. This comprehensive copyright relevance test of product descriptions and images provides a more comprehensive assessment of whether a target product details page infringes copyright.
[0037] Step S5400: When the product details page is judged to be a trademark infringement or copyright infringement, a corresponding infringement detection report is constructed and pushed to the user, ending the listing process of the target product.
[0038] When trademark infringement or copyright infringement is detected on the product details page through the previous steps, the report generation mechanism will be immediately activated. The generation of the infringement detection report is a systematic process that aims to provide users with comprehensive and detailed infringement information. In one embodiment, the generated infringement detection report will clearly indicate the type of infringement, i.e. trademark infringement or copyright infringement, as well as the specific content of the infringement. For example, if it is a trademark infringement, the infringement detection report will detail which words in the product description text are highly similar to the preset trademark text information, or which logos in the product image match the trademark brand logo. If it is a copyright infringement, the infringement detection report will list the similarity between the product description text and the copyright description text set, as well as the matching between the product image and the copyright image set. This information will help users quickly understand the specific manifestations of the infringement.
[0039] In some embodiments, in addition to infringement information and evidence, the infringement detection report will also include relevant legal tips and suggestions. Based on the type of infringement and the specific circumstances, it will provide references to corresponding legal clauses to help users understand the legal consequences of infringement. At the same time, the report also provides recommended handling measures, such as suggesting that merchants immediately modify product description text and remove infringing words or logos; or suggesting that merchants redesign product images to avoid using unauthorized copyrighted images, so as to help users correct infringements in a timely manner and reduce possible legal risks.
[0040] After an infringement report is generated, it is pushed to users through various channels. For merchants on e-commerce platforms, notifications can be sent directly through the merchant management backend, displaying the infringement detection report to the merchant in the form of a pop-up window or message. In some embodiments, email or SMS notifications are also sent to ensure that merchants receive the report in a timely manner. For platform operators, the report is uploaded to the e-commerce platform's infringement management module for further review and processing.
[0041] When the infringement detection report is pushed, the listing process of the target product will be automatically terminated to prevent the infringing products from continuing to be listed and sold, avoiding causing greater damage to the rights holder. Merchants can only continue to list the product after modifying the product information according to the report content and passing the platform's re-review.
[0042] In one embodiment, the e-commerce platform provides an appeal mechanism. If a merchant disagrees with the infringement detection results, they can submit an appeal through the appeal channel. The e-commerce platform will then forward the appeal information to a professional infringement review team, who will conduct a manual review of the infringement. During the review process, the reviewer will re-evaluate the infringement evidence and make a final judgment based on the supplementary information provided by the merchant. If the review results show that the product is not infringing, the product's listing restrictions will be lifted and a corresponding notification will be sent to the merchant. If the review results uphold the original judgment, the merchant will be notified of the review results and the product's infringement status will continue to be maintained.
[0043] Through the above steps, not only can infringement detection reports be generated quickly and accurately, but also effective communication and timely processing of information can be ensured. This mechanism not only helps protect intellectual property rights, but also provides merchants with clear operational guidance, allowing them to correct infringements in a timely manner and maintain a healthy operating environment for e-commerce platforms.
[0044] It can be seen from the typical embodiments of the present application that the technical solution of the present application has many advantages, including but not limited to the following aspects: This application, by pre-positioning an infringement detection mechanism in the product listing process, can promptly detect and prevent potential infringements before the products are officially listed, thus avoiding the discovery of infringing products only after they have been widely disseminated on the platform, thereby minimizing damage to trademark and copyright owners, while also maintaining the business reputation and legal compliance of the e-commerce platform.
[0045] Specifically, this application determines whether there is trademark infringement by judging whether there are word segments in the product description text whose similarity with the preset trademark text information exceeds a threshold, and whether the product image contains the preset trademark brand logo. At the same time, by confirming the correlation between the product description text and product image and the preset copyright description text set and copyright image set, it determines whether there is copyright infringement. When the product details page is judged to be a trademark infringement or copyright infringement, an infringement detection report is constructed and pushed to the user, and the product listing process is terminated at the same time. That is, this application realizes comprehensive infringement detection of product details pages, which not only covers the product description text, but also includes product images. By comprehensively analyzing the infringement information in the text and images, it can more comprehensively identify trademark infringement and copyright infringement. This multi-dimensional detection method significantly improves the accuracy and comprehensiveness of infringement detection, and avoids infringements that may be missed by a single-dimensional detection method.
[0046] Based on any embodiment of the method of the present application, determining whether the product description text contains a segmentation whose similarity to the preset trademark text information is greater than a preset similarity threshold, and if so, determining that the product details page constitutes trademark infringement, includes: Step S5210: performing word segmentation processing on the product description text to obtain multiple target word segments, wherein the product description text includes a product title, product details, and brand information; Obtain the product description text of the target product. These product description texts cover multiple aspects of the product, including product title, product details, and brand information. The product title is a short description carefully designed by the merchant to attract consumers' attention, which includes the core features of the product or the brand name; the product details are more detailed, including the function, use, material, specifications, etc. of the product; the brand information is a specific introduction or statement of the product brand, and the product description text is segmented. For example, for the product title "Brand A Sports Shoes", after word segmentation, vocabulary units such as "Brand A", "Sports", and "Shoes" will be obtained. The accuracy of word segmentation directly affects the reliability of subsequent similarity calculations. Therefore, by adopting advanced word segmentation algorithms and combining them with professional dictionaries in the product field, the accuracy and completeness of the word segmentation results can be ensured.
[0047] Step S5220: Calculate a first similarity between the target segmentation corresponding to the product title and the trademark text information, a second similarity between the target segmentation corresponding to the product details information and the trademark text information, and a third similarity between the target segmentation corresponding to the brand information and the trademark text information; After completing word segmentation of product descriptions, the results are categorized by text source, corresponding to product titles, product details, and brand information. Different parts of the text have varying importance and sensitivity in trademark infringement detection. For example, product titles are more directly appealing to consumers and are therefore more sensitive to trademark usage; brand information, on the other hand, is directly linked to trademark usage and is a key area for infringement detection.
[0048] In one embodiment, for the product title, the target word segmentation corresponding to the product title is compared one by one with the trademark text information. Since product titles are usually short and contain concentrated information, a combination of exact matching and fuzzy matching is used. Exact matching refers to directly comparing whether the word segmentation and the trademark text are exactly the same, while fuzzy matching determines whether there are variants or similar words by calculating the similarity, and calculates the corresponding first similarity. For the product details information, the target word segmentation is also extracted and the similarity is calculated with the trademark text information. The product details information is longer and contains richer descriptive content, so the semantic similarity algorithm is used to process it. The semantic similarity algorithm considers not only the surface form of the word, but also the semantic meaning of the word. For example, through word embedding technology (such as Word2Vec or BERT), the word segmentation and trademark text can be converted into semantic vectors, and the cosine similarity between these vectors is calculated as the second similarity. For brand information, we pay special attention to the use of brand names and their variants. Brand information is directly related to the use of trademarks. We directly extract target words from the brand information and perform exact and fuzzy matching with the trademark text information to obtain the third similarity. Due to the high sensitivity of brand names, a stricter similarity threshold is used, such as 0.9.
[0049] By calculating the similarity between the target word and trademark text information in the product title, product details, and brand information respectively, we can more comprehensively assess the trademark infringement risk in the product description text.
[0050] Step S5230: When the first similarity is greater than a preset first similarity threshold, the second similarity is greater than a preset second similarity threshold, or the third similarity is greater than a preset third similarity threshold, it is determined that there is trademark infringement on the target product details page.
[0051] After calculating the similarity between the target word segmentation for the product title, product details, and brand information and the trademark text in the previous step, these similarity values are compared with the corresponding preset thresholds. These thresholds are set based on the actual needs and experience of trademark infringement detection and are used to distinguish between normal text descriptions and infringements. For example, the first similarity threshold for product titles is set at 0.8, the second similarity threshold for product details is set at 0.7, and the third similarity threshold for brand information is set at 0.9. These thresholds take into account the importance and sensitivity of different text components in trademark infringement detection.
[0052] When the first similarity between the target word in the product title and the trademark text exceeds a preset first similarity threshold, the product title is judged to be at risk of trademark infringement. Similarly, when the second similarity between the target word in the product details and the trademark text exceeds a preset second similarity threshold, the product details are judged to be at risk of trademark infringement. Product details contain richer descriptive content, so its similarity threshold is set relatively low to capture a wider range of potential infringements. Similarly, when the third similarity between the target word in the brand information and the trademark text exceeds a preset third similarity threshold, the brand information is judged to be at risk of trademark infringement. The similarity between the product title, product details, and brand information is comprehensively considered. As long as the similarity of any of these components exceeds the corresponding preset threshold, the target product details page is judged to be at risk of trademark infringement. This comprehensive judgment method ensures the comprehensiveness and accuracy of trademark infringement detection, avoiding missed detections or false detections due to misjudgment of a single text component.
[0053] In this embodiment, by segmenting product descriptions and calculating similarity for each, trademark infringement can be more accurately detected, avoiding the potential misjudgments or omissions that can result from single-dimensional detection. Secondly, different similarity thresholds are set for product titles, product details, and brand information. This differentiated approach fully accounts for the varying sensitivity of different text components in trademark infringement detection, further improving detection accuracy.
[0054] Based on any embodiment of the method of the present application, determining whether the product image contains a preset trademark brand logo, and if so, determining that the product details page constitutes trademark infringement, includes: Step S5240: Input the product image into a logo detection model to obtain a detection result of whether a preset trademark brand logo exists in the product image and a corresponding confidence level; After obtaining product images of the target product, these images are fed into a pre-trained logo detection model. This logo detection model, built using deep learning technology, can automatically identify trademarks and brand logos in images. The model training phase will be discussed in detail later and will not be detailed here.
[0055] During the detection process, the logo detection model analyzes the input product image pixel by pixel, extracts key features from the product image, and compares them with the preset trademark and brand logo. The logo detection model then outputs a detection result that clearly indicates whether the trademark and brand logo are present in the image. The model also provides a confidence score, which reflects the model's confidence in the detection result. The confidence score is typically a value between 0 and 1, with higher values indicating higher confidence in the detection result. For example, if the target product image contains a trademark logo, the logo detection model will output a detection result after analyzing the product image, indicating the presence of the trademark logo and providing a confidence score, such as 0.98, indicating that the logo in the image is highly consistent with the preset trademark and brand logo.
[0056] In one embodiment, the logo detection model employs a variety of advanced technologies to improve detection accuracy and robustness. For example, the logo detection model incorporates a convolutional neural network (CNN) to extract local features from an image and utilizes fully connected layers to integrate these features, thereby achieving accurate recognition of trademark logos. Furthermore, the logo detection model employs data augmentation techniques, rotating, scaling, and cropping training samples to enhance the model's ability to recognize trademark logos of varying shapes and sizes.
[0057] By inputting product images into the logo detection model and providing corresponding confidence levels for subsequent confidence level judgment, trademark brand logos in product images can be identified efficiently and accurately, thereby providing strong technical support for trademark infringement detection.
[0058] Step S5250: When the detection result indicates that a preset trademark brand logo exists in the product image and the confidence level is greater than or equal to a preset confidence level threshold, the product details page is determined to be trademark infringing. When the detection result shows the presence of a preset trademark brand logo in a product image, the confidence level of the detection result is further checked to see if it meets the preset confidence threshold. The preset confidence threshold is set by those skilled in the art based on the actual application scenario and the requirements for detection accuracy. Assuming the preset confidence threshold is 0.95, after the logo detection model analyzes a product image, the output detection result shows the presence of the trademark logo of the brand "Brand A" in the image, and the corresponding confidence level is 0.98. Since this confidence level is higher than the preset confidence threshold, it is directly determined that trademark infringement has occurred in the product image.
[0059] In some embodiments, batch detection is performed on multiple product images. For example, if a product details page contains multiple images, each image is analyzed by the logo detection model, and the detection results and confidence levels are evaluated individually. If the detection results for any of the images meet the aforementioned conditions, i.e., the presence of a trademark logo and a confidence level greater than or equal to a preset threshold, the product details page is determined to constitute trademark infringement.
[0060] This confidence-based judgment mechanism can effectively distinguish between high-confidence infringements and possible misjudgments, thereby protecting the legitimate rights and interests of trademark owners while avoiding unnecessary troubles for merchants.
[0061] Step S5260: When the detection result indicates that a preset trademark brand logo exists in the product image and the confidence level is less than the confidence level threshold, determine whether the text information in the product image contains the trademark text information; if so, determine that the product details page is a trademark infringement.
[0062] After analyzing a product image, if the logo detection model outputs a detection result indicating the presence of a pre-set trademark or brand logo, but the corresponding confidence level falls below the preset confidence threshold, further verification begins. This may be due to unclear logos, interference, or differences in the logo's form from the training examples, leading to low confidence in the model's output. Further verification of the text within the product image can improve the accuracy of the judgment.
[0063] Textual information in product images exists in various forms, such as watermarks, labels, and advertising slogans. This textual information may include trademarks, brand names, or other trademark-related content. In some embodiments, optical character recognition (OCR) technology is used to extract textual information from images and compare it with pre-set trademark textual information. OCR technology can convert the text in the image into an editable text format, facilitating subsequent text analysis. Suppose a product image contains a less-than-clear trademark logo, and the logo detection model's confidence level is 0.85, which is lower than the pre-set confidence threshold of 0.90. In this case, OCR technology is activated to extract text from the image. If the OCR technology identifies the trademarked brand name "Brand A" in the image, the similarity between the extracted text and the pre-set trademark textual information is further calculated. For example, by calculating the edit distance or cosine similarity, it can be determined whether the textual content closely matches the trademark textual information. If the similarity between the OCR-extracted textual information and the trademark textual information exceeds a pre-set textual similarity threshold (e.g., 0.8), the textual information in the product image is determined to contain trademark textual information, and the product details page is therefore deemed to infringe trademark rights. This dual verification mechanism, which combines image logo detection and text information analysis, can effectively make up for the shortcomings of a single detection method and improve the accuracy and reliability of infringement detection.
[0064] In this embodiment, it is possible to effectively handle infringement detection issues in low-confidence situations, avoiding misjudgments or missed judgments due to unclear identification or interference factors. By analyzing the text information in the image through OCR technology, the existence of the trademark is further confirmed, thereby protecting the legitimate rights and interests of trademark owners while reducing misjudgments of merchants. This dual verification method that combines visual recognition and text analysis not only improves the accuracy of detection, but also enhances the robustness of the infringing product detection method on e-commerce platforms, enabling it to adapt to more complex scenarios.
[0065] Based on any embodiment of the method of the present application, before inputting the product image into the logo detection model and obtaining a detection result of whether a preset trademark brand logo exists in the product image and a corresponding confidence level, the method includes: Step S6100: Acquire a sample image set containing a preset trademark brand logo, and label the sample image set with the trademark brand logo; Collect a large number of sample images containing pre-set trademark and brand logos to train a detection model that can accurately identify trademark and brand logos. These sample images must contain a variety of pre-set trademark and brand logos to ensure that the model can learn the characteristics of different trademark logos. Sample images can come from a variety of sources, including but not limited to official brand images, authorized advertising images, and common product packaging images in the market. These images should cover a variety of scenes, angles, lighting conditions, and backgrounds to enhance the model's generalization capabilities.
[0066] After collecting sample images, they are further annotated in detail. The purpose of annotation is to clearly indicate the specific location, shape, size, brand name, and other information of the trademark brand logo in each image. For example, if the sample image contains the trademark logo of the brand "Brand A", the annotator needs to mark the bounding box of the logo in the image and indicate the brand name as "Brand A". This annotation process usually requires professional annotation tools to ensure the accuracy and consistency of the annotation. It should be noted that during the annotation process, detailed annotation specifications need to be formulated to avoid annotation inconsistencies caused by subjective differences among annotators. For example, the annotation specifications specify how the bounding box is drawn and the annotation format of the brand name.
[0067] In addition to annotating the location and name of the trademark brand logo, other features of the trademark logo can also be further annotated, such as color, font, degree of deformation, etc. This additional annotation information can help the logo detection model better learn the visual characteristics of the trademark logo, thereby improving detection accuracy.
[0068] Step S6200: Based on the labeled sample image set, train the logo detection model to a convergent state, so that the logo detection model learns the ability to detect whether there is a trademark brand logo in the input product image.
[0069] After collecting and labeling a set of sample images containing trademark and brand logos in the previous step, these labeled images are used to train a logo detection model. The logo detection model is built on deep learning techniques, such as convolutional neural networks (CNNs), which excel in image recognition tasks and can automatically extract features from images. At the beginning of the training process, the logo detection model randomly initializes its parameters. The logo detection model then receives batches of labeled sample images and attempts to predict the location and category of the trademark and brand logo in each image. Since these sample images are already labeled, the logo detection model can evaluate its performance by comparing its predictions with the labeled information. This comparison is achieved using a loss function, which calculates the difference between the predicted results and the actual annotations and outputs a numerical value representing the magnitude of this difference. For example, if the logo detection model's predicted trademark logo location differs significantly from the actual labeled location, the loss function will result in a higher value.
[0070] After each prediction, the logo detection model adjusts its parameters based on the loss function to reduce prediction error, a process known as backpropagation. By continuously adjusting its parameters, the logo detection model gradually learns to more accurately identify trademarks and brand logos. During training, the logo detection model undergoes multiple iterations, each of which involves multiple passes through the entire sample image set. As training progresses, the logo detection model's parameters are gradually optimized, and the loss function value gradually decreases, indicating that the logo detection model's predictions are increasingly close to the actual annotation information.
[0071] During the training process, various techniques are employed to prevent model overfitting. Overfitting occurs when a model performs well on the training set but poorly on new, unseen data. To prevent overfitting, data augmentation techniques are used to generate more training data by rotating, scaling, cropping, and color-adjusting training samples, thereby enhancing the model's generalization capabilities. Furthermore, regularization techniques, such as L2 regularization, are used to limit the complexity of model parameters by adding a regularization term to the loss function, thereby reducing the risk of overfitting.
[0072] When the logo detection model reaches a certain level of training, it enters a state of convergence. This state means that the model's parameters change very little in successive iterations, or the loss function no longer decreases significantly. At this point, the logo detection model has learned to detect the presence of brand logos in input product images.
[0073] This embodiment effectively improves the accuracy and generalization of trademark and brand logo detection. Through comprehensive sample image annotation and deep learning model training, the model not only identifies the location and category of trademark logos but also learns their various visual features, enabling accurate recognition in diverse scenarios, angles, lighting conditions, and backgrounds. Furthermore, through data augmentation and regularization techniques, the model's generalization capabilities are significantly enhanced, enabling it to better adapt to unseen data and thus demonstrate greater robustness in practical applications.
[0074] Based on any embodiment of the method of the present application, determining whether the product image contains a preset trademark brand logo, and if so, determining that the product details page constitutes trademark infringement, includes: Step S7100: Segment the product image according to preset size specifications to obtain multiple identification primitives; Before image segmentation, the segmentation granularity is determined based on preset size specifications. These size specifications are based on common brand logo sizes and the actual image resolution to ensure that the segmented logo elements contain sufficient detail for subsequent detection. For example, if the brand logo is small, a smaller segmentation size is selected; if the brand logo is large, a larger segmentation size is selected.
[0075] After obtaining the product image, the image is evenly divided according to the preset size specifications. This process is implemented through image processing algorithms, such as using sliding window technology. The sliding window will move pixel by pixel on the product image according to the set size. Each movement will capture an image area of the same size as the window, thereby generating multiple identification primitives.
[0076] The segmented logo primitives will be extracted one by one for subsequent trademark and brand logo detection. These logo primitives contain various parts of the product, such as the product itself, packaging, background, etc. By segmenting the image into multiple primitives, each area in the image can be analyzed more carefully, thus avoiding the omission of trademark logos due to their small size or unclear location.
[0077] In some embodiments, some optimization algorithms are used to improve the efficiency of image segmentation. For example, if the background of a product image is relatively simple, the outline of the product is first located through an edge detection algorithm, and then the product area is segmented, thereby reducing unnecessary calculations. This method can not only improve the segmentation speed, but also improve the accuracy of detection because the background area usually does not contain trademark logos.
[0078] Step S7200: performing feature matching on the multiple logo primitives and the trademark and brand logos in the logo database, and screening out logo primitives whose similarity exceeds a preset fourth similarity threshold; After image segmentation in the previous step, multiple logo primitives are generated. Each primitive represents a localized area of the product image. These logo primitives need to be feature-matched with trademark and brand logos in a logo database to determine whether they contain trademark and brand logos. The logo database is a collection of images containing registered trademark and brand logos. These logo images have been carefully processed and annotated to serve as a reference standard for detection.
[0079] Before feature matching, image features are extracted for each logo element. Image features include color distribution, texture features, and shape features. For example, a color histogram can be used to describe the color distribution of the element, a texture analysis algorithm can be used to extract texture features, or an edge detection algorithm can be used to extract shape features. These features can describe the image content from different perspectives, providing rich information for subsequent matching. Simultaneously, the same features are extracted for the trademark and brand logos in the logo database, so that each trademark and brand logo will also have a corresponding set of feature vectors.
[0080] The features of the logo graphic elements are compared one by one with the features of the trademark brand logos in the logo database. The comparison process is achieved by calculating the similarity between the feature vectors. For example, algorithms such as Euclidean distance, Hamming distance or cosine similarity are used to measure the similarity between two feature vectors, and the logo graphic elements whose similarity exceeds the preset fourth similarity threshold are screened out. The preset fourth similarity threshold is set according to the actual application scenario and the requirements for detection accuracy. A higher threshold can reduce false alarms, but may miss some infringing logos with lower similarity; a lower threshold can increase the sensitivity of detection, but may increase the false alarm rate. Therefore, it is necessary to reasonably set the similarity threshold based on the characteristics of the trademark logo and actual needs.
[0081] Step S7300: Determine the image center distances between each pair of the filtered identification graphic elements in the product image, and calculate the average of the center distances; To further confirm whether the identified graphics filtered out in the previous step constitute infringement, it is necessary to analyze the spatial position relationship of these graphics in the image. Specifically, determine the image center position of each logo. The image center position is obtained through simple geometric calculations and is the center point of the logo bounding box. For example, if the position of a logo in the image is defined by a rectangular frame, the center point of the rectangular frame is used as the image center of the graphics. In this way, a clear center position can be determined for each filtered logo.
[0082] Then, the image center distance between each pair of markers is calculated. This image center distance is calculated based on the coordinates of the image center, that is, the straight-line distance between the center points of each pair of markers is measured to determine the distance between all the selected markers.
[0083] After calculating all pairwise image center distances, we average these distances. This is calculated by adding all distances and dividing by the number of distances. By averaging these distances, we can assess the distribution of the selected logo elements within the product images. A smaller average distance indicates that these logo elements are more concentrated within the image, while a larger average distance indicates that these logo elements are more dispersed within the image.
[0084] Step S7400: When the average center distance is less than a preset distance threshold, it is determined that the product details page constitutes trademark infringement.
[0085] This step compares the average center distance with a preset distance threshold. The preset distance threshold is set by a person skilled in the art based on the typical layout and actual application scenarios of the trademark logo, and is used to distinguish between normal trademark use and possible infringement. When the similarity between these logo elements and the trademark brand logo has exceeded the preset fourth similarity threshold, and the average center distance of these logo elements in the product image is less than the preset distance threshold, it indicates that these logo elements are relatively concentrated. When the user combines two authorized trademark logos to circumvent infringement detection, this step can be used to identify and determine this potential infringement, that is, because the user deliberately combines or juxtaposes multiple trademark logos to confuse the audience and circumvent infringement detection. For example, the user places two similar trademark logos very close to each other, or disperses the various parts of the trademark logo but keeps them in a small space, in an attempt to evade infringement detection of a single logo in this way. The judgment mechanism of this step can effectively identify situations where detection is circumvented by abnormal combinations or layouts of trademark logos, thereby ensuring the comprehensiveness and accuracy of trademark infringement detection.
[0086] In this embodiment, it is possible to effectively handle trademark infringement issues in complex scenarios, especially when infringers attempt to evade detection by combining or juxtaposing multiple trademark logos. By combining image segmentation and spatial distribution analysis, this application can not only identify infringements of a single logo, but also detect abnormal combinations of multiple logos.
[0087] Based on any embodiment of the method of the present application, determining the correlation between the product description text and product image and a preset set of copyright description text and copyright image, and determining that the product details page constitutes copyright infringement when any of the correlations meets a preset condition, includes: Step S5310: Based on the product description text, retrieve the text feature vector with the highest similarity from a first vector database containing a set of copyright description texts, and determine the corresponding first similarity; In one embodiment, an efficient vector retrieval system is constructed to store and retrieve feature vectors of copyright description text. The core of the vector retrieval system is a first vector database, which can quickly process large-scale vector data and support efficient similarity retrieval. In this step, FAISS (Facebook AI SimilaritySearch) can be used as the vector retrieval system. FAISS is an open source library designed for large-scale vector retrieval and clustering. It supports multiple distance metrics, such as cosine similarity and Euclidean distance. It can quickly find the vector most similar to the target vector in massive data, that is, it can quickly retrieve the text feature vector with the highest similarity in the first vector database. In order to ensure the accuracy of the retrieval, the copyright description text needs to be preprocessed and feature extracted. Preprocessing includes steps such as text cleaning and word segmentation. For example, the copyright description text can be segmented using the Jieba word segmentation tool. After preprocessing, the copyright description text needs to be feature extracted to convert the text into a numerical vector that can describe its semantics. In this embodiment, the BERT (Bidirectional Encoder Representations from Transformers) model is used for feature extraction. BERT is a pre-trained language model based on the Transformer architecture that can capture bidirectional semantic information in the text. By inputting the copyright description text into the BERT model, the feature vectors of each text can be obtained. These feature vectors can reflect the semantic content of the text, such as the theme, emotion, style, etc.
[0088] In some embodiments, the BERT model is fine-tuned to further improve the quality of the text feature vector. Fine-tuning refers to further training the BERT model using data from a specific field based on the pre-trained model to better adapt it to the current task. In this step, the copyright description text is used as fine-tuning data to train the BERT model so that it can better capture the semantic features of the copyright text.
[0089] The product description text of the target product also needs to be preprocessed and feature extracted to finally obtain the corresponding target feature vector. Then, the obtained target feature vector is searched for similarity with the copyright description text feature vector in the first vector database. The vector retrieval system can quickly calculate the similarity between the target feature vector and all text feature vectors in the database, and return the most similar text feature vector and its similarity value, i.e., the first similarity value. In some embodiments, cosine similarity is used as a similarity measurement method. The value range of cosine similarity is between [-1, 1]. The closer the value is to 1, the more similar the two vectors are.
[0090] Through this step, we can accurately determine whether the product details page constitutes copyright infringement based on the semantic similarity between the product description text and the copyright description text. This not only enables the rapid processing of large-scale text data, but also effectively captures the semantic information of the text, providing an efficient and accurate technical means for infringement detection on product details pages of e-commerce platforms.
[0091] Step S5320: Based on the product image, retrieve the image feature vector with the highest similarity from a second vector database containing a set of copyrighted images, and determine the corresponding second similarity; In one embodiment, a vector retrieval system for storing and retrieving copyright image feature vectors is constructed. The core of the vector retrieval system is a second vector database. The constructed second vector database can process large-scale image feature data and support efficient similarity retrieval. In this embodiment, FAISS (Facebook AI Similarity Search) is used as the vector retrieval system. Before building the vector retrieval system, the copyright image set needs to be preprocessed and feature extracted. The preprocessing includes adjusting the image resolution, cropping, normalization, and other operations to ensure that all images have consistent format and size. For example, all copyright images are uniformly adjusted to a size of 224×224 pixels so that the subsequent feature extraction model can process them. At the same time, the images are normalized, that is, the pixel value range is adjusted to between [0, 1] or [-1, 1] to reduce the impact of the numerical range on feature extraction.
[0092] After preprocessing, feature extraction is performed on the copyrighted image to convert the image into a numerical vector that can describe its visual content. In one embodiment, a pre-trained convolutional neural network (CNN) model is used for feature extraction. For example, models such as ResNet, VGG, or CLIP can be used. These models have been pre-trained on large-scale image datasets and can extract image features with semantic meaning. Taking the CLIP model as an example, the CLIP model is a multimodal model that can process image and text data simultaneously and extract semantic associations between images and text. When extracting image features, the copyrighted image is input into the CLIP model, and the CLIP model will output a fixed-length feature vector that can capture the visual content and semantic information of the image, such as the shape, color, texture, and scene of the object in the image.
[0093] The target product image also requires preprocessing and feature extraction to obtain its feature vector. This feature vector is then compared with the feature vectors of copyrighted images in the second vector database for similarity. FAISS can quickly calculate the similarity between the target vector and all vectors in the database and return the most similar vector and its similarity value.
[0094] Step S5330: When the first similarity is greater than a preset text similarity threshold or the second similarity is greater than a preset image similarity threshold, the product details page is determined to be copyright infringing.
[0095] A text similarity threshold and an image similarity threshold are pre-set. These are key parameters for determining whether a product details page constitutes copyright infringement and are set by those skilled in the art based on actual application scenarios and copyright protection requirements. During infringement detection, both text similarity and image similarity must be considered comprehensively. Specifically, when the first similarity between the product description text and the copyright description text is greater than a preset text similarity threshold, or when the second similarity between the product image and the copyright image is greater than a preset image similarity threshold, the product details page can be determined to constitute copyright infringement. This comprehensive assessment method can more comprehensively cover various manifestations of infringement and avoid misjudgments or omissions based on a single dimension.
[0096] In this embodiment, copyright infringement detection can be fully covered in both text and image dimensions. By combining the semantic similarity of text and the visual similarity of images, an efficient and accurate method for detecting infringing goods is provided. This embodiment can not only quickly process large-scale text and image data, but also effectively capture the semantic information of text and images, thereby demonstrating higher accuracy and reliability in protecting copyright.
[0097] Based on any embodiment of the method of the present application, based on the product description text, before retrieving the text feature vector with the highest similarity from a first vector database containing a set of copyright description texts and determining the corresponding first similarity, the method includes: Step S8100: Obtain a preset copyright description text set, perform word segmentation and semantic vectorization on each copyright description text in the copyright description text set, generate a corresponding text feature vector, associate the copyright description text with the copyright description text, and store it in the first vector database; In one embodiment, a corpus containing copyright description text is constructed. This corpus covers various copyright text content related to the target product. This copyright text content can be obtained from channels such as documents provided by the copyright holder, official websites, and authorized e-commerce platform pages. The obtained text data needs to be cleaned and preprocessed to remove irrelevant symbols, stop words, and repeated content to ensure the quality and consistency of the corpus. The copyright description text is then segmented. After the segmentation is completed, each segmented text unit is semantically vectorized using pre-trained language models based on deep learning, such as BERT and RoBERTa. These models can capture the semantic information of the text through unsupervised learning on large-scale corpora. Specifically, the segmented text is input into the pre-trained model to obtain a context-sensitive embedding vector for each word. Then, through aggregation operations (such as averaging and max pooling), the vectors of multiple words are combined into a feature vector for the entire text. In addition, traditional methods such as TF-IDF can be combined to weight the vectors to highlight the contribution of important words.
[0098] The generated text feature vectors must be associated with the corresponding copyright description text and stored in a first vector database. This database utilizes efficient vector retrieval systems such as FAISS (Facebook AI Similarity Search) or Milvus. These systems can quickly perform similarity searches within large-scale vector data sets and support a variety of distance metrics, such as cosine similarity and Euclidean distance. During storage, a unique identifier must be assigned to each text feature vector, and metadata such as the source and category of the corresponding copyright description text must be recorded to facilitate rapid location and retrieval during subsequent searches.
[0099] This step builds a high-quality, efficient database of copyright description text vectors, providing a solid foundation for subsequent infringement detection based on product descriptions. During infringement detection, when determining whether a product description constitutes copyright infringement, the text feature vector with the highest similarity to the product description can be directly retrieved from this vector database. In this way, infringement can be determined based on a similarity threshold.
[0100] Step S8200: Obtain a preset copyright image set, perform image feature extraction on each copyright image in the copyright image set, generate a corresponding image feature vector, associate the corresponding copyright image, and store it in the second vector database.
[0101] The copyright image collection contains all copyrighted images related to the target product. These copyrighted images come from official image libraries provided by the copyright holder, authorized e-commerce platform pages, brand promotional materials, etc. The collected copyrighted images should cover different angles, backgrounds, and lighting conditions to ensure the comprehensiveness and robustness of feature extraction. After collecting the copyrighted image collection, each image needs to be preprocessed, and then image features are extracted from the preprocessed copyrighted images. In this step, deep learning models can be used to extract image features. For example, pretrained convolutional neural network (CNN) models such as ResNet, VGG, or CLIP can be used. These models have been pretrained on large-scale image datasets and can extract semantically meaningful image features.
[0102] The generated image feature vectors need to be associated with the corresponding copyrighted images and stored in a second vector database. This database can utilize efficient vector retrieval systems such as FAISS or Milvus, which can quickly perform similarity searches within large-scale vector data and support multiple distance metrics. During storage, a unique identifier must be assigned to each image feature vector, along with metadata such as the source and category of the corresponding copyrighted image, to facilitate rapid location and retrieval during subsequent searches.
[0103] Through this step, a high-quality and efficient copyright image feature vector database is constructed. When it is necessary to determine whether a product image constitutes copyright infringement, the image feature vector with the highest similarity to the product image can be directly retrieved from the vector database, and whether there is infringement can be determined based on the similarity threshold.
[0104] In this embodiment, two different types of copyrighted content, text and images, can be processed simultaneously, a high-quality and efficient copyright feature vector database is constructed, and fast similarity retrieval is achieved through an efficient vector retrieval system (such as FAISS or Milvus). It supports multiple distance measurement methods, thereby quickly determining whether product description text and product images constitute copyright infringement.
[0105] See also Figure 3According to one aspect of the present application, an infringing product detection device is provided, comprising an event response module 5100, a trademark infringement detection module 5200, a copyright infringement detection module 5300, and a result processing module 5400, wherein the event response module 5100 is configured to respond to a target product listing event triggered by a user, initiate a pre-infringement detection process in the listing process of the target product, and obtain the product description text and product image in the product details page of the target product; the trademark infringement detection module 5200 is configured to determine whether the product description text contains a similarity greater than a preset trademark text information. The similarity threshold is used for word segmentation, or whether the product image contains a preset trademark brand logo. When it does, the product details page is determined to constitute trademark infringement; the copyright infringement detection module 5300 is configured to confirm the correlation between the product description text and product image and the preset copyright description text set and copyright image set. When any correlation meets the preset conditions, the product details page is determined to constitute copyright infringement; the result processing module 5400 is configured to construct a corresponding infringement detection report and push it to the user when the product details page is determined to be a trademark infringement or copyright infringement, thereby ending the listing process of the target product.
[0106] Based on any embodiment of the device of the present application, the trademark infringement detection module 5200 includes: a text segmentation submodule, configured to perform segmentation processing on the product description text to obtain multiple target segmentations, wherein the product description text includes a product title, product details information and brand information; a similarity calculation submodule, configured to respectively calculate a first similarity between the target segmentation corresponding to the product title and the trademark text information, a second similarity between the target segmentation corresponding to the product details information and the trademark text information, and a third similarity between the target segmentation corresponding to the brand information and the trademark text information; a trademark infringement determination submodule, configured to determine that there is trademark infringement on the target product details page when the first similarity is greater than a preset first similarity threshold, the second similarity is greater than a preset second similarity threshold, or the third similarity is greater than a preset third similarity threshold.
[0107] Based on any embodiment of the device of the present application, the trademark infringement detection module 5200 further includes: an image identification detection submodule, configured to input the product image into the identification detection model to obtain a detection result of whether a preset trademark brand identification exists in the product image and the corresponding confidence level; a high-confidence judgment submodule, configured to determine that the product details page is a trademark infringement when the detection result indicates that a preset trademark brand identification exists in the product image and the confidence level is greater than or equal to a preset confidence threshold; a low-confidence verification submodule, configured to determine whether the text information in the product image contains the trademark text information when the detection result indicates that a preset trademark brand identification exists in the product image and the confidence level is less than the confidence threshold, and if so, determine that the product details page is a trademark infringement.
[0108] On the basis of any embodiment of the device of the present application, the image logo detection submodule includes: an image set annotation submodule, which is configured to obtain a sample image set containing a preset trademark brand logo, and annotate the sample image set with the trademark brand logo; a model training submodule, which is configured to train the logo detection model to a convergence state based on the annotated sample image set, so that the logo detection model learns the ability to detect whether there is a trademark brand logo in the input product image.
[0109] Based on any embodiment of the device of the present application, the trademark infringement detection module 5200 further includes: an image segmentation processing submodule, configured to perform image segmentation on the product image according to preset size specifications to obtain multiple identification graphics elements; a trademark feature matching submodule, configured to perform feature matching on the multiple identification graphics elements with the trademark brand logos in the identification database, and screen out identification graphics elements whose similarity exceeds a preset fourth similarity threshold; a spatial distribution analysis submodule, configured to determine the image center distance between each of the screened identification graphics elements in the product image, and calculate the average center distance; a spatial distribution determination submodule, configured to determine that the product details page constitutes trademark infringement when the average center distance is less than a preset distance threshold.
[0110] Based on any embodiment of the device of the present application, the copyright infringement detection module 5300 includes: a text feature retrieval submodule, which is configured to retrieve the text feature vector with the highest similarity from a first vector database containing a set of copyright description texts based on the product description text, and determine the corresponding first similarity; an image feature retrieval submodule, which is configured to retrieve the image feature vector with the highest similarity from a second vector database containing a set of copyright pictures based on the product picture, and determine the corresponding second similarity; a copyright infringement determination submodule, which is configured to determine that the product details page is copyright infringing when the first similarity is greater than a preset text similarity threshold or the second similarity is greater than a preset image similarity threshold.
[0111] On the basis of any embodiment of the device of the present application, the text feature retrieval submodule includes: a text feature library construction submodule, which is configured to obtain a preset copyright description text set, perform word segmentation processing and semantic vectorization processing on each copyright description text in the copyright description text set, generate a corresponding text feature vector, associate the copyright description text with it, and store it in the first vector database; an image feature library construction submodule, which is configured to obtain a preset copyright picture set, perform image feature extraction processing on each copyright picture in the copyright picture set, generate a corresponding image feature vector, associate the corresponding copyright picture, and store it in the second vector database.
[0112] In order to solve the above technical problems, the embodiment of the present application also provides a computer device. Figure 4 As shown, a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database may store a control information sequence. When the computer-readable instructions are executed by the processor, the processor may implement a method for detecting infringing goods. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor may execute the infringing goods detection method of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0113] In this embodiment, the processor is used to execute Figure 3The memory stores the program code and various data required to execute the modules and submodules in the device. The network interface is used to transmit data between user terminals or servers. The memory in this embodiment stores the program code and data required to execute all modules and submodules in the infringing goods detection device of this application. The server can call the server's program code and data to execute the functions of all submodules.
[0114] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the infringing product detection method of any embodiment of the present application.
[0115] Those skilled in the art will appreciate that all or part of the processes in the above-described embodiments of the present application can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes in the above-described embodiments of the method. The aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0116] Those skilled in the art will understand that the various operations, methods, steps, measures, and schemes in the processes discussed in this application may be interchanged, changed, combined, or deleted. Furthermore, other steps, measures, and schemes in the various operations, methods, and processes discussed in this application may also be interchanged, changed, rearranged, decomposed, combined, or deleted. Furthermore, the steps, measures, and schemes in the various operations, methods, and processes in the prior art that are open source and disclosed in this application may also be interchanged, changed, rearranged, decomposed, combined, or deleted.
[0117] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for detecting infringing goods, characterized in that: include: In response to a user-triggered target product listing event, the pre-infringement detection process in the listing process of the target product is initiated, and the product description text and product image on the product details page of the target product are obtained; Determine whether the product description text contains a word whose similarity to the preset trademark text information is greater than a preset similarity threshold, or whether the product image contains a preset trademark brand logo. If so, determine that the product details page constitutes trademark infringement; Determining the relevance between the product description text and product image and a preset set of copyright description text and copyright image. If any of the relevances meets a preset condition, the product details page is determined to constitute copyright infringement. When the product details page is judged to be a trademark infringement or copyright infringement, a corresponding infringement detection report is constructed and pushed to the user, ending the listing process of the target product.
2. The method for detecting infringing goods according to claim 1, characterized in that: Determining whether the product description text contains a word whose similarity to the preset trademark text information is greater than a preset similarity threshold, and if so, determining that the product details page constitutes trademark infringement, including: Performing word segmentation processing on the product description text to obtain multiple target word segments, wherein the product description text includes a product title, product details information, and brand information; Calculating respectively a first similarity between the target segmentation word corresponding to the product title and the trademark text information, a second similarity between the target segmentation word corresponding to the product details information and the trademark text information, and a third similarity between the target segmentation word corresponding to the brand information and the trademark text information; When the first similarity is greater than a preset first similarity threshold, the second similarity is greater than a preset second similarity threshold, or the third similarity is greater than a preset third similarity threshold, it is determined that there is trademark infringement on the target product details page.
3. The method for detecting infringing goods according to claim 1, characterized in that: Determine whether the product image contains a preset trademark brand logo. If so, determine that the product details page constitutes trademark infringement, including: Input the product image into the logo detection model to obtain a detection result of whether a preset trademark brand logo exists in the product image and the corresponding confidence level; When the detection result indicates that a preset trademark brand logo exists in the product image and the confidence level is greater than or equal to a preset confidence level threshold, the product details page is determined to be trademark infringing; When the detection result indicates that there is a preset trademark brand logo in the product image and the confidence level is less than the confidence level threshold, it is determined whether the text information in the product image contains the trademark text information. If so, the product details page is determined to be a trademark infringement.
4. The method for detecting infringing goods according to claim 3, characterized in that: Before inputting the product image into the logo detection model to obtain a detection result of whether a preset trademark brand logo exists in the product image and a corresponding confidence level, the method includes: Obtaining a sample image set containing a preset trademark brand logo, and labeling the sample image set with the trademark brand logo; Based on the labeled sample image set, the logo detection model is trained to convergence, so that the logo detection model learns the ability to detect whether there is a trademark brand logo in the input product image.
5. The method for detecting infringing goods according to claim 1, characterized in that: Determine whether the product image contains a preset trademark brand logo. If so, determine that the product details page constitutes trademark infringement, including: Segment the product image according to preset size specifications to obtain multiple identification primitives; Performing feature matching on the multiple logo elements with the trademark and brand logos in the logo database, and screening out logo elements whose similarity exceeds a preset fourth similarity threshold; Determine the image center distances between each of the filtered identification graphic elements in the product image, and calculate the average of the center distances; When the average center distance is less than a preset distance threshold, it is determined that the product details page constitutes trademark infringement.
6. The method for detecting infringing goods according to claim 1, characterized in that: Determine the relevance between the product description text and product image and a preset set of copyright description text and copyright image. If any of the relevances meets a preset condition, determine that the product details page constitutes copyright infringement, including: Based on the product description text, retrieve the text feature vector with the highest similarity from a first vector database containing a set of copyright description texts, and determine the corresponding first similarity; Based on the product image, searching for the image feature vector with the highest similarity from a second vector database containing a set of copyrighted images, and determining a corresponding second similarity; When the first similarity is greater than a preset text similarity threshold or the second similarity is greater than a preset image similarity threshold, the product details page is determined to be copyright infringing.
7. The method for detecting infringing goods according to claim 6, characterized in that: Based on the product description text, before retrieving the text feature vector with the highest similarity from a first vector database containing a set of copyright description texts and determining the corresponding first similarity, the method includes: Obtaining a preset copyright description text set, performing word segmentation and semantic vectorization on each copyright description text in the copyright description text set, generating a corresponding text feature vector, associating the copyright description text with the first vector database; A preset copyright image set is obtained, image feature extraction is performed on each copyright image in the copyright image set, a corresponding image feature vector is generated, and the associated copyright image is stored in the second vector database.
8. A device for detecting infringing goods, characterized in that: include: An event response module is configured to respond to a user-triggered target product listing event, initiate a pre-infringement detection process in the target product listing process, and obtain product description text and product images from the target product's product details page; a trademark infringement detection module configured to determine whether the product description text contains a word whose similarity to a preset trademark text information is greater than a preset similarity threshold, or whether the product image contains a preset trademark brand logo, and if so, determine that the product details page constitutes trademark infringement; a copyright infringement detection module configured to determine the relevance between the product description text and product image and a preset set of copyright description text and copyright image, and to determine that the product details page constitutes copyright infringement when any of the relevances meets a preset condition; The result processing module is configured to construct a corresponding infringement detection report and push it to the user when the product details page is judged to be a trademark infringement or copyright infringement, thereby ending the listing process of the target product.
9. An infringing goods detection device, comprising a central processing unit and a memory, characterized in that: The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.
10. A non-volatile readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.
Citation Information
Patent Citations
Infringement trademark detection method based on Logo identification
CN115393748A
Commodity infringement detection method and device, equipment, medium and product
CN115620019A
Commodity page detection system and detection method thereof
CN117709970A
Identification method, computing device, computer storage medium and computer program product
CN117910450A
Trademark infringement detection method and system
CN118691857A
Cited By
Trademark infringement detection method and system based on commodity pictures
CN121962872A
A trademark infringement detection method and system based on commodity pictures
CN121962872B