Infringement commodity detection method and device, equipment and medium
By introducing a pre-infringement detection mechanism into the product listing process on e-commerce platforms and combining it with text and image analysis technology, the problem of low efficiency and insufficient accuracy in detecting infringements on product detail pages on e-commerce platforms has been solved, achieving comprehensive identification and timely prevention of trademark and copyright infringements.
Patent Information
- Application Number
- CN202510948720.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-07-10
AI Technical Summary
In existing technologies, infringement detection on product detail pages on e-commerce platforms is inefficient and inaccurate, manual review is easily affected by subjective factors, and single-dimensional technical means are unable to fully identify complex infringements, leading to misjudgments or missed judgments.
By introducing a pre-infringement detection mechanism in the product listing process, comprehensively analyzing product description text and images, and using natural language processing and image recognition technology to determine trademark and copyright infringement, including similarity comparison and relevance detection, an infringement detection report is generated and the listing process is terminated.
It achieves comprehensive infringement detection on product detail pages, improves the accuracy and comprehensiveness of detection, timely discovers and prevents potential infringements, protects intellectual property rights, and maintains the business reputation and legal compliance of e-commerce platforms.
Smart Images

Figure CN120450828B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of e-commerce, and in particular to a method and device for detecting infringing goods, equipment and a medium. BACKGROUND
[0002] Under the background of rapid development of e-commerce, infringement in the goods detail page has gradually become a problem to be solved. With the popularity of e-commerce platforms, the display and dissemination of goods information have become more convenient, but at the same time, trademark infringement and copyright infringement have occurred frequently. These infringements not only damage the legitimate rights and interests of brand owners and original creators, but also damage the fair competition environment of the market.
[0003] Traditional infringing goods detection methods mainly rely on manual review or single-dimensional technical means. Although the manual review method can accurately judge the infringement, it is difficult to meet the rapid review needs of a large amount of goods information due to its low efficiency, and the manual review is also easily affected by subjective factors, resulting in inconsistency of the review results. On the other hand, single-dimensional technical means, such as simple string matching or rule-based image recognition, can improve the efficiency of infringement detection to some extent, but their accuracy and comprehensiveness still need to be improved. For example, the simple string matching method can only detect the same text content, and it is powerless for infringement with similar semantics but different expressions, and the rule-based image recognition method is difficult to deal with complex image backgrounds and various trademark styles, and is prone to misjudgment or omission. SUMMARY
[0004] The primary purpose of the present application is to solve at least one of the above problems and provide a method and device for detecting infringing goods, equipment and a medium.
[0005] According to one aspect of the present application, a method for detecting infringing goods is provided, comprising the following steps:
[0006] In response to a target goods listing event triggered by a user, starting a pre-infringement detection process in the listing process of the target goods, obtaining the goods description text and goods pictures in the goods detail page of the target goods;
[0007] Determining whether the goods description text contains a word divided between the preset trademark text information with a similarity greater than a preset similarity threshold, or whether the goods pictures contain a preset trademark brand logo. When the goods detail page contains a trademark infringement, it is determined that the goods detail page constitutes a trademark infringement;
[0008] Confirming the relevance between the goods description text and the goods pictures and the preset copyright description text set and the copyright picture set. When any relevance meets a preset condition, it is determined that the goods detail page constitutes a copyright infringement.
[0009] When the product detail page is determined to be trademark infringement or copyright infringement, a corresponding infringement detection report is constructed and pushed to the user, and the listing process of the target product is ended.
[0010] According to another aspect of the present application, a kind of infringement product detection device is provided, comprising:
[0011] Event response module is set to respond to the target product listing event triggered by user, start the pre-infringement detection process in the listing process of the target product, obtain the product description text and product picture in the product detail page of target product;
[0012] Trademark infringement detection module is set to determine whether the product description text contains the word segmentation between the preset trademark text information Similarity greater than the preset similarity threshold, or whether the product picture contains the preset trademark brand logo, when containing, determine that the product detail page constitutes trademark infringement;
[0013] Copyright infringement detection module is set to confirm the correlation between the product description text and product picture and the preset copyright description text set and copyright picture set, when any correlation meets the preset condition, determine that the product detail page constitutes copyright infringement;
[0014] Result processing module is set to when the product detail page is determined to be trademark infringement or copyright infringement, corresponding infringement detection report is constructed and pushed to the user, and the listing process of the target product is ended.
[0015] According to another aspect of the present application, a kind of infringement product detection equipment is provided, including central processing unit and memory, the central processing unit is used to call the computer program stored in the memory to execute the steps of the infringement product detection method described in the present application.
[0016] According to another aspect of the present application, a kind of non-volatile readable storage medium is stored with the computer program realized according to the infringement product detection method in the form of computer readable instruction, when the computer program is called and runs by computer, the steps included in the method are executed.
[0017] The technical scheme of the present application has many advantages, including but not limited to the following aspects:
[0018] The present application can discover and prevent potential infringement behavior in time before the product is formally listed by pre-infringement detection mechanism in product listing process, avoid the infringement product to be discovered after widespread dissemination on platform, thereby maximumly reduce the damage to trademark and copyright right holder, and also maintain the commercial credibility and legal compliance of e-commerce platform.
[0019] Specifically, the present application determines whether there is trademark infringement by judging whether there are words in the commodity description text that have a similarity exceeding a threshold with the preset trademark text information, and whether the commodity picture contains the preset trademark brand logo, and determines whether there is copyright infringement by confirming the relevance between the commodity description text and the commodity picture and the preset copyright description text set and copyright picture set, and when the commodity detail page is determined to be trademark infringement or copyright infringement, constructs an infringement detection report and pushes it to the user, and at the same time ends the listing process of the commodity. That is, the present application realizes comprehensive infringement detection of the commodity detail page, which not only covers the commodity description text, but also includes the commodity picture, and through comprehensive analysis of the infringement information in the text and the picture, it can more comprehensively identify trademark infringement and copyright infringement behavior. This multi-dimensional detection method significantly improves the accuracy and comprehensiveness of infringement detection, and avoids the infringement behavior that may be missed by single-dimensional detection method. BRIEF DESCRIPTION OF DRAWINGS
[0020] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description, taken in conjunction with the following drawings, in which:
[0021] Figure 1 a network architecture of an exemplary e-commerce platform of the present application;
[0022] Figure 2 a flowchart of a typical embodiment of the infringement commodity detection method of the present application;
[0023] Figure 3 a principle block diagram of the infringement commodity detection device of the present application;
[0024] Figure 4 a structural schematic diagram of an infringement commodity detection device adopted by the present application. DETAILED DESCRIPTION
[0025] The embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as a limitation of the present application.
[0026] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0027] Those skilled in the art will understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless specifically defined as such, will not be interpreted in an idealized or overly formal sense.
[0028] like Figure 1 In the network architecture shown, the e-commerce platform 82 is deployed on the Internet to provide corresponding services to its users. Similarly, the devices 80 of the merchant users of the e-commerce platform 82 and the devices 81 of the consumer users are also connected to the Internet to use the services provided by the e-commerce platform.
[0029] The exemplary e-commerce platform 82 provides supply and demand matching of products and / or services to the general public with the help of Internet infrastructure. In the e-commerce platform 82, products and / or services are provided as commodity information. To simplify the description, the concepts of commodity, product, etc. are used in this application to refer to the products and / or services in the e-commerce platform 82, which may specifically be physical products, digital products, tickets, service subscriptions, other offline services, etc.
[0030] Real-world entities can access the e-commerce platform 82 in the identity of users to use various online services provided by the e-commerce platform 82 to achieve the purpose of participating in the business activities implemented by the e-commerce platform 82. These entities can be natural persons, legal persons, or social organizations, etc. Corresponding to the two types of entities of merchants and consumers in the business activities, the e-commerce platform 82 correspondingly exists two types of users of merchant users and consumer users. The entities of the product circulation chain in the business activities, including manufacturers, sellers, retailers, logistics providers, etc., can use online services in the e-commerce platform 82 in the identity of merchant users, while the consumers in the business activities, including real or potential consumers, can use online services in the e-commerce platform 82 in the identity of their corresponding consumer users. In actual business activities, the same entity can act as a merchant user and a consumer user, and flexible and variable understanding should be made accordingly.
[0031] The infrastructure for deploying the e-commerce platform 82 mainly includes a backend architecture and front-end devices. The backend architecture runs various online services through a service cluster, including middleware or front-end services for platform parties, services for consumers, services for merchants, etc., to enrich and perfect its service functions; the front-end devices mainly cover terminal devices used by users as clients to access the e-commerce platform 82, including but not limited to various mobile terminals, personal computers, point-of-sale devices, etc. For example, a merchant user can use his terminal device 80 to input product information for his online store, or use the interface provided by the e-commerce platform to generate product information; a consumer user can access the web page of the online store implemented by the e-commerce platform 82 through his terminal device 81, trigger the shopping process through the shopping button provided on the web page, and call various online services provided by the e-commerce platform 82 in the shopping process, so as to achieve the purpose of shopping and ordering.
[0032] In some embodiments, the e-commerce platform 82 can be implemented by a processing facility including a processor and a memory, which stores a set of instructions that, when executed, cause the e-commerce platform 82 to perform the electronic commerce and support functions involved in this application. The processing facility can be part of a server, a client, a network infrastructure, a mobile computing platform, a cloud computing platform, a fixed computing platform, or other computing platforms, and provides electronic components of the e-commerce platform 82, merchant devices, payment gateways, application developers, marketing channels, transportation providers, customer devices, point-of-sale devices, etc.
[0033] The e-commerce platform 82 can be implemented as a cloud computing service, a software as a service (SaaS), an infrastructure as a service (IaaS), a platform as a service (PaaS), a desktop as a service (DaaS), a hosted software as a service, a mobile backend as a service (MBaaS), an information technology management as a service (ITMaaS), and the like online services. In some embodiments, various functional components of the e-commerce platform 82 can be implemented to be suitable for operation on various platforms and operating systems, for example, an administrator user of an online store enjoys the same or similar functions regardless of various embodiments such as iOS, Android, HomonyOS, or web pages.
[0034] The e-commerce platform 82 can implement respective independent stations for various merchants to run their respective online stores, and provide respective instances of the business management engine for the merchants to establish, maintain, and run one or more online stores in one or more independent stations. The instance of the business management engine can be used for content management, task automation, and data management of one or more online stores, and can configure various specific business processes of the online store through an interface or a built-in component to support the implementation of business activities. The independent station is the infrastructure of the e-commerce platform 82 with cross-border service functions, and the merchant can maintain its online store based on the independent station in a relatively centralized and autonomous manner. The independent station usually has a domain name and storage space dedicated to the merchant, and different independent stations have relative independence. The e-commerce platform 82 can provide standardized or personalized technical support for a large number of independent stations, so that the merchant user can customize a business management engine instance suitable for itself, and use this business management engine instance to maintain one or more online stores owned by the merchant.
[0035] The online store can be configured and maintained by the merchant user logging in as an administrator to the business management engine instance. With the support of various online services provided by the infrastructure of the e-commerce platform 82, the merchant user can configure various functions in the online store and review various data as an administrator, for example, the merchant user can manage various aspects of the online store, such as viewing recent activities of the online store, updating the product catalog of the online store, managing orders, recent visit activities, total order activities, and the like; the merchant user can also view more detailed information about the business and visitors to the online store of the merchant by obtaining reports or metrics, such as displaying sales summaries of the overall business of the merchant, specific sales and engagement data of the marketing channels of the activity sales, and the like.
[0036] The e-commerce platform 82 can provide communication facilities and associated merchant interfaces for providing electronic communications and marketing, such as with electronic message aggregation facilities to collect and analyze communication interactions between merchants, consumers, merchant devices, customer devices, point-of-sale devices, etc., to aggregate and analyze communications, such as to increase the potential for providing product sales, etc. For example, a consumer can have a question about a product that can generate a conversation between the consumer and a merchant (or an automated processor-based agent representing the merchant), where the communication facilities are responsible for the interaction and provide the merchant with analysis on how to improve the probability of a sale.
[0037] In some embodiments, an application suitable for installation to a terminal device can be provided to serve the access needs of different users, so as to enable various users to access the e-commerce platform 82, such as a merchant back-end module of an online store in the e-commerce platform 82, etc., in the terminal device by running the application. In the process of implementing business activities through these functions, the e-commerce platform 82 can implement various functions related to the implementation of business activities as middleware or online services and open corresponding interfaces, and then implant the interface access functions into the application to implement the function extension and task implementation. The business management engine can include a series of basic functions, and expose these functions to online services and / or application calls through APIs. The online services and the application use the corresponding functions by remotely calling the corresponding APIs.
[0038] With the support of various components of the commerce management engine instance, the e-commerce platform 82 can provide online shopping functions, enabling merchants to establish contact with customers in a flexible and transparent manner, and consumer users can select and purchase items online, create product orders, provide delivery addresses for goods in product orders, and complete payment confirmation of product orders. Then, the merchant can review and complete or cancel the order. The review component carried by the commerce management engine instance can enable the compliant use of business processes to ensure that the order is suitable for fulfillment before actual fulfillment. Orders can sometimes be fraudulent and need to be verified, there is a payment method that requires merchants to wait to ensure that funds are received, and so on. Order risk can be generated by fraud detection tools submitted by third parties through order risk APIs, etc. Before fulfillment, the merchant can obtain payment information or wait to receive payment information in order to mark the order as paid, and only then can the merchant prepare to deliver the product. Similar situations can be reviewed accordingly. The review process can be implemented by the fulfillment component. The merchant can review, adjust work, and trigger related fulfillment services with the help of the fulfillment component, such as: manual fulfillment services when the merchant selects and packages products in a box, purchases shipping labels and enters their tracking numbers, or simply marks the item as fulfilled; custom fulfillment services that can define sending emails for notifications; API fulfillment services that can trigger third-party applications to create fulfillment records on third parties; legacy fulfillment services that can trigger custom API calls from the commerce management engine to third parties; gift card fulfillment services. It can provide generating numbers and activating gift cards. Merchants can use order printer applications to print shipping labels. The fulfillment process can be performed when the item is packaged in a box and ready for shipping, tracking, delivery, consumer receipt verification, etc.
[0039] As can be seen, the services provided by the e-commerce platform are based on products as the core, and the corresponding product data is the basic data of the e-commerce platform. The product information is provided through the product data, and the mining and utilization of the product data are the basis for realizing various technical services, including using user transaction data and product data in the product data of the e-commerce platform to provide basic services for the operation of the data processing system. Therefore, the data processing system can run in any one or more servers of the e-commerce platform cluster, so as to realize various functions by utilizing various product data provided by the e-commerce platform.
[0040] Referring to Figure 2 In a typical embodiment of the infringement product detection method of the present application, the method comprises the following steps:
[0041] Step S5100, in response to a target product listing event triggered by a user, starting a pre-infringement detection process in the listing process of the target product, obtaining product description text and product pictures in the product detail page of the target product;
[0042] In the operation process of the e-commerce platform, the merchant needs to upload the product information to the product detail page of the e-commerce platform, so that consumers can browse and purchase. However, with the increasing number of products in the e-commerce platform, the infringement behavior in the product detail page also gradually increases. In order to effectively prevent infringement in the product detail page and ensure the healthy operation of the e-commerce platform, the present application will start the pre-infringement detection process to detect the infringement of the product detail page corresponding to the product to be listed before the product is listed. That is, when the merchant triggers the listing event of the target product on the e-commerce platform, the infringement product detection method of the present application is automatically triggered.
[0043] Specifically, first, the product description text and product pictures in the product detail page of the target product are obtained. The product description text is a detailed description of the product by the merchant, including product title, product function, use method, specification parameter, and brand information, etc. important information. The product description text is one of the important bases for judging trademark infringement. Since trademark infringement often involves the unauthorized use of others' trademark text or logo, at this time, if the natural language processing technology is used to analyze the obtained product description text, extract the key information therein, and compare it with the preset trademark text information, it can be determined whether there is a possibility of trademark infringement. The product picture is an important content to show the appearance of the target product, including the real picture, the packaging picture, the use scene picture, etc. of the product. The picture infringement detection mainly aims at the unauthorized use of others' copyright pictures, or the unauthorized use of others' trademark logo in the product picture. At this time, if the image recognition technology is used to analyze the product picture, extract the key features in the picture, and compare them with the preset copyright picture set and trademark logo, it can be determined whether there is an infringement behavior.
[0044] Through this pre-infringement detection mechanism, the e-commerce platform can discover potential infringement behavior in the early stage of product listing and take timely measures such as preventing product listing, notifying the merchant to modify the information, etc., so as to effectively protect the intellectual property rights and maintain the commercial order of the platform.
[0045] Step S5200, judging whether the product description text contains a word segmentation with a similarity greater than a preset similarity threshold with the preset trademark text information, or whether the product picture contains a preset trademark logo. When it contains, it is determined that the product detail page constitutes trademark infringement.
[0046] First, the text of the commodity description is analyzed in depth, including word segmentation, which means breaking down sentences and paragraphs in the commodity description text into individual word units. Trademark infringement often involves the unauthorized use of key words or phrases of others' trademarks, and word segmentation can help identify these potential infringing words more accurately. For example, if the description of the target commodity contains the brand name "Brand A", which has been registered as a trademark, "Brand A" will be identified.
[0047] The segmented words are compared with the preset trademark text information. The commodity text information is stored in the commodity text information library, and the trademark text information library is a database containing registered trademark keywords. The similarity of each word in the commodity description text to the words in the trademark text information library is calculated. Similarity calculation can use various algorithms, such as cosine similarity algorithm, which measures the similarity between two text vectors by calculating the cosine of the angle between them. If the similarity of a word to a trademark in the trademark information library exceeds the preset similarity threshold, it will be marked as a potential infringing word. For example, if the preset similarity threshold is 0.8, and the similarity of the segmented words corresponding to the commodity description text to the "Brand A" trademark in the trademark information library is 0.9, it is determined that there is trademark infringement in the commodity description text.
[0048] In image recognition, the commodity picture is detected for trademark brand identification. In one embodiment, a sample image set containing trademark brand identification is needed, which is carefully annotated, including information such as the location, shape and features of the trademark. Based on these annotated sample images, a model is trained to identify whether the preset trademark brand identification exists in the commodity picture. For example, if the picture of the target commodity contains the "Brand A" hook-shaped trademark identification, the model will detect the existence of the hook-shaped trademark identification. When the identification detection model detects the existence of the trademark brand identification in the commodity picture, the confidence of the identification is further determined. Confidence is an indicator of the reliability of the model's detection results, a value between 0 and 1. If the confidence is greater than or equal to the preset confidence threshold, for example 0.95, it is directly determined that there is trademark infringement in the commodity picture. However, if the confidence is lower than the preset threshold, the text information in the commodity picture needs to be further verified to see if it contains trademark text information. Considering that the trademark identification in the commodity picture may not be clear enough or there may be some interference, resulting in a low confidence of the model detection, the combination of text information verification can more accurately determine whether there is trademark infringement.
[0049] Through the above comprehensive analysis of the product description text and the product picture, it can be comprehensively judged whether the target product detail page constitutes trademark infringement. This multi-dimensional detection method not only improves the accuracy of infringement detection, but also effectively deals with various complex infringement behaviors.
[0050] Step S5300, confirming the relevance between the product description text and the product picture and the preset copyright description text set and the copyright picture set, when any relevance meets the preset condition, determining that the product detail page constitutes copyright infringement;
[0051] First, the copyright relevance analysis is performed on the product description text. The copyright description text set is a database containing text content that has been protected by copyright, including product titles, product details, or brand information in various product detail pages. The product description text of the target product is compared with the copyright description text set to determine whether there is unauthorized similar or identical content. That is, the product description text is compared with the corresponding copyright description text in the copyright description text set. In the comparison process, natural language processing technology is used to perform semantic vectorization processing on the product description text and the copyright description text. Semantic vectorization is a technology that converts text content into numerical vectors, which can capture the semantic features of the text and thus realize quantitative analysis of the similarity of the text. The similarity between the product description text and the copyright description text is calculated, and the preset condition is set to be that if the similarity exceeds the preset text similarity threshold, for example 0.7, it is determined that there is a high correlation between the product description text and the copyright description text, and it is determined that the product detail page constitutes copyright infringement.
[0052] Further, the copyright relevance detection is performed on the product picture. The copyright picture set is a database containing picture content that has been protected by copyright. The copyright pictures in this database can be composed of various product pictures uploaded from various product detail pages. The product picture of the target product is compared with the copyright picture set to determine whether there is an unauthorized similar or identical picture. In the picture comparison process, image feature extraction technology is used to extract the feature vectors of the product picture and the copyright picture. Image feature extraction technology can capture the visual features of the picture, such as color distribution, texture, shape, etc. The similarity between the product picture and the copyright picture is calculated, for example by calculating the Euclidean distance or Hamming distance between the two image feature vectors. The preset condition is set to be that if the similarity exceeds the preset image similarity threshold, for example 0.8, it is determined that there is a high correlation between the product picture and the copyright picture, and it is determined that the product detail page constitutes copyright infringement.
[0053] The relevance of the commodity description text and the commodity picture is comprehensively considered. If the relevance between any content in the commodity description text or the commodity picture and the copyright description text set or the copyright picture set meets a preset condition, it is determined that the commodity detail page constitutes copyright infringement. Through the above comprehensive copyright relevance detection of the commodity description text and the commodity picture, whether the target commodity detail page constitutes copyright infringement can be more comprehensively judged.
[0054] Step S5400, when the commodity detail page is determined to be trademark infringement or copyright infringement, a corresponding infringement detection report is constructed and pushed to the user, and the listing process of the target commodity is ended.
[0055] When the commodity detail page is detected to have trademark infringement or copyright infringement behavior through the previous steps, a report generation mechanism is immediately started. The generation of the infringement detection report is a systematic process, which aims to provide comprehensive and detailed infringement information for the user. In an embodiment, the generated infringement detection report will explicitly indicate the type of infringement, i.e., trademark infringement or copyright infringement, and the specific infringement content. For example, if it is trademark infringement, the infringement detection report will specify which words in the commodity description text are highly similar to the preset trademark text information, or which logo in the commodity picture matches the trademark brand logo. If it is copyright infringement, the infringement detection report will list the similarity between the commodity description text and the copyright description text set, and the matching between the commodity picture and the copyright picture set. These information will help the user quickly understand the specific performance of the infringement.
[0056] In some embodiments, in addition to the infringement information and evidence, the infringement detection report will also contain relevant legal tips and suggestions. According to the type of infringement and the specific circumstances, the corresponding legal provisions are provided to help the user understand the legal consequences of the infringement behavior. At the same time, the report also gives the recommended handling measures, such as recommending the merchant to immediately modify the commodity description text, remove the infringing words or logo; or recommending the merchant to redesign the commodity picture to avoid using unauthorized copyright pictures, so as to help the user correct the infringement behavior in time and reduce the possible legal risk.
[0057] After the generation of the infringement report, the report is pushed to the user through various channels. For the merchants on the e-commerce platform, the merchant management background can be used to directly send a notification, and the infringement detection report can be displayed to the merchant in the form of a pop-up window or a message. In some embodiments, an email or an SMS message is also sent to ensure that the merchant can receive the report in time. For platform operators, the report is uploaded to the infringement management module of the e-commerce platform for further review and processing by the operators.
[0058] At the same time of pushing the infringement detection report, the listing process of the target commodity is automatically terminated to prevent the infringing commodity from being continuously listed for sale, so as to avoid causing greater damage to the right holder, and the merchant can only continue to list the commodity after modifying the commodity information according to the report content and passing the re-audit of the platform.
[0059] In an embodiment, the e-commerce platform provides an appeal mechanism, if the merchant has objections to the infringement detection result, the merchant can submit an appeal application through the appeal channel, the e-commerce platform transfers the appeal information to a professional infringement audit team, and an artificial audit personnel reviews the infringement, in the review process, the audit personnel re-evaluates the infringement evidence, combines the supplementary information provided by the merchant, and makes a final judgment. If the review result shows that the commodity does not infringe, the listing restriction of the commodity is lifted, and a corresponding notification is sent to the merchant; if the review result maintains the original judgment, the merchant is notified of the review result, and the infringement state of the commodity is continued to be maintained.
[0060] Through the above steps, not only can the infringement detection report be quickly and accurately generated, but also the effective transmission and timely processing of information can be ensured. This mechanism not only helps to protect intellectual property rights, but also provides clear operation guidelines for merchants to timely correct infringement behavior and maintain a healthy operating environment for e-commerce platforms.
[0061] According to the typical embodiments of the present application, it can be known that the technical solutions of the present application have many advantages, including but not limited to the following aspects:
[0062] The present application can discover and prevent potential infringement behavior in time before the commodity is formally listed by pre-positioning the infringement detection mechanism in the commodity listing process, avoiding the discovery of infringing commodities after they are widely spread on the platform, thereby minimizing the damage to trademark and copyright right holders, and at the same time maintaining the commercial credibility and legal compliance of the e-commerce platform.
[0063] Specifically, the present application determines whether there is trademark infringement by judging whether there are words in the commodity description text that have a similarity exceeding a threshold with the preset trademark text information, and whether the commodity picture contains the preset trademark brand logo, and determines whether there is copyright infringement by confirming the relevance between the commodity description text and the commodity picture and the preset copyright description text set and copyright picture set. When the commodity detail page is determined to be trademark infringement or copyright infringement, an infringement detection report is constructed and pushed to the user, and at the same time the listing process of the commodity is ended. That is, the present application realizes comprehensive infringement detection of the commodity detail page, which not only covers the commodity description text but also includes the commodity picture. By comprehensively analyzing the infringement information in the text and the picture, trademark infringement and copyright infringement behavior can be more comprehensively identified. This multi-dimensional detection method significantly improves the accuracy and comprehensiveness of infringement detection, avoiding the infringement behavior that may be missed by a single-dimensional detection method.
[0064] On the basis of any embodiment of the method of the present application, it is judged whether the commodity description text contains a word segmentation with a similarity greater than a preset similarity threshold to the preset trademark text information. When it contains, it is determined that the commodity detail page constitutes trademark infringement, comprising:
[0065] Step S5210, performing word segmentation processing on the commodity description text to obtain a plurality of target word segments, the commodity description text including commodity title, commodity detail information and brand information;
[0066] Obtain the commodity description text of the target commodity, which covers multiple aspects of the commodity, including commodity title, commodity detail information and brand information, etc. The commodity title is a short description carefully designed by the merchant to attract the attention of consumers, which contains the core features or brand name of the commodity; the commodity detail information is more detailed, including the function, purpose, material, specification, etc. of the commodity; the brand information is a specific introduction or statement of the commodity brand. For example, for the commodity title "Brand A sports shoes", the word segmentation processing will get "Brand A", "sports", "shoes" and other vocabulary units. The accuracy of word segmentation directly affects the reliability of subsequent similarity calculation, so by using advanced word segmentation algorithm combined with professional dictionary in the commodity field, the accuracy and integrity of the word segmentation result can be ensured.
[0067] Step S5220, respectively calculating the first similarity between the target word segment corresponding to the commodity title and the trademark text information, the second similarity between the target word segment corresponding to the commodity detail information and the trademark text information, and the third similarity between the target word segment corresponding to the brand information and the trademark text information;
[0068] After completing the word segmentation processing of the commodity description text, the word segmentation results are classified according to the text source, corresponding to the commodity title, commodity detail information and brand information respectively. The importance and sensitivity of different parts of the text in trademark infringement detection are different. For example, the commodity title is more directly attracting the attention of consumers, so the use of trademarks is more sensitive; while the brand information is directly related to the use of trademarks, which is the key area of infringement detection.
[0069] In one embodiment, for the product title part, the target word pieces corresponding to the product title are compared one by one with the trademark text information. Since the product title is usually short and information-intensive, a combination of exact matching and fuzzy matching is used. Exact matching refers to directly comparing whether the word pieces and the trademark text are exactly the same, while fuzzy matching judges whether there are variants or approximate words by calculating the similarity, and the corresponding first similarity is calculated. For the product detail information part, the target word pieces are also extracted and compared with the trademark text information, and the product detail information is longer and contains more descriptive content, so a semantic similarity algorithm is used to process it. The semantic similarity algorithm not only considers the surface form of the words, but also considers the semantic meaning of the words. For example, through word embedding technology (such as Word2Vec or BERT), the word pieces and the trademark text can be converted into semantic vectors, and the cosine similarity between these vectors is calculated as the second similarity. For the brand information part, special attention is paid to the use of brand names and their variants, and the brand information is directly related to the use of trademarks. The target word pieces in the brand information are directly extracted and compared with the trademark text information using exact matching and fuzzy matching to obtain the third similarity. Since brand names are highly sensitive, a stricter similarity threshold is used, for example, 0.9.
[0070] By calculating the similarity between the target word pieces in the product title, product detail information, and brand information and the trademark text information respectively, the risk of trademark infringement in the product description text can be more comprehensively evaluated.
[0071] Step S5230, when the first similarity is greater than the preset first similarity threshold, the second similarity is greater than the preset second similarity threshold, or the third similarity is greater than the preset third similarity threshold, it is determined that the target product detail page exists trademark infringement.
[0072] After the similarity calculation of the target word pieces in the product title, product detail information, and brand information and the trademark text information is completed in the previous step, these similarity values are compared with the corresponding preset thresholds. These thresholds are set according to the actual needs and experience of trademark infringement detection, and are used to distinguish the boundary between normal text description and infringement. For example, the first similarity threshold of the product title is set to 0.8, the second similarity threshold of the product detail information is 0.7, and the third similarity threshold of the brand information is 0.9. These threshold settings take into account the importance and sensitivity of different text parts in trademark infringement detection.
[0073] When the first similarity between the target word in the commodity title and the trademark text information is greater than the preset first similarity threshold, it is determined that the commodity title part has a trademark infringement risk. Similarly, when the second similarity between the target word in the commodity detail information and the trademark text information is greater than the preset second similarity threshold, it is determined that the commodity detail information part has a trademark infringement. The commodity detail information contains more descriptive content, and therefore the similarity threshold is set relatively low, so as to more widely capture potential infringement. Similarly, when the third similarity between the target word in the brand information and the trademark text information is greater than the preset third similarity threshold, it is determined that the brand information part has a trademark infringement. The similarity in the commodity title, the commodity detail information and the brand information is comprehensively considered. As long as the similarity of any one part exceeds the corresponding preset threshold, it is determined that the target commodity detail page has a trademark infringement. This comprehensive judgment method can ensure the comprehensiveness and accuracy of trademark infringement detection, and avoid missing or false detection due to false judgment of a single text part.
[0074] In the embodiment, by performing word segmentation processing on the commodity description text and calculating the similarity respectively, the trademark infringement behavior can be more accurately captured, and false judgment or omission caused by single dimension detection can be avoided. Secondly, different similarity thresholds are set for the commodity title, the commodity detail information and the brand information. This differentiated processing method fully considers the sensitivity difference of different text parts in trademark infringement detection, and further improves the accuracy of detection.
[0075] On the basis of any embodiment of the method of the application, it is judged whether the commodity picture contains a preset trademark brand logo. When it contains, it is determined that the commodity detail page constitutes a trademark infringement, comprising:
[0076] In step S5240, the commodity picture is input into an identification detection model to obtain a detection result of whether the preset trademark brand logo exists in the commodity picture and a corresponding confidence degree.
[0077] After obtaining the commodity pictures of the target commodity, the commodity pictures are input into a pre-trained identification detection model. The identification detection model is constructed based on deep learning technology and can automatically identify the trademark brand logo in the picture. The model training stage is described in the subsequent specific embodiments, which will not be described here.
[0078] In the detection process, the identification detection model performs pixel-by-pixel analysis on the input product picture, extracts key features in the product picture, and compares them with the preset trademark brand identification. Then the identification detection model outputs a detection result, which clearly indicates whether there is a trademark brand identification in the picture. At the same time, the model gives a confidence value, which reflects the model's confidence in the detection result. The confidence value is usually a value between 0 and 1. The higher the value, the higher the confidence of the model in the detection result. For example, suppose the target product picture contains a trademark identification. After analyzing the product picture, the identification detection model outputs a detection result indicating that there is a trademark identification in the picture and gives a confidence value, such as 0.98. At this time, it is determined that the identification in the picture matches the preset trademark brand identification.
[0079] In an embodiment, the identification detection model uses various advanced techniques to improve detection accuracy and robustness. For example, the identification detection model combines convolutional neural networks (CNN) to extract local features of the picture, and uses fully connected layers to integrate these features, thereby achieving accurate identification of trademark identification. In addition, the identification detection model also uses data augmentation techniques to enhance the identification detection model's ability to recognize trademark identification of different shapes and sizes by rotating, scaling, cropping, and other operations on training samples.
[0080] By inputting the product picture into the identification detection model and giving the corresponding confidence value for subsequent confidence judgment, the trademark brand identification in the product picture can be efficiently and accurately identified, thereby providing strong technical support for trademark infringement detection.
[0081] Step S5250, when the detection result represents that the preset trademark brand identification exists in the product picture and the confidence is greater than or equal to the preset confidence threshold, it is determined that the product detail page is trademark infringement;
[0082] When the detection result shows that the product picture contains the preset trademark brand identification, further check whether the confidence of the detection result meets the preset confidence threshold. The preset confidence threshold is set by a person skilled in the art according to the actual application scenario and the requirement for detection accuracy. Suppose the preset confidence threshold is 0.95, and the identification detection model analyzes a product picture and outputs a detection result showing that there is a trademark identification of "Brand A" in the picture, and the corresponding confidence is 0.98. Since the confidence is higher than the preset confidence threshold, it is directly determined that there is a trademark infringement in the product picture.
[0083] In some embodiments, batch detection is performed on multiple product pictures. For example, a product detail page contains multiple pictures, each of which needs to be analyzed by the logo detection model, and the detection result and confidence of each picture are evaluated one by one. If any of the pictures meets the above conditions, i.e., there is a trademark logo and the confidence is greater than or equal to the preset threshold, it is determined that the product detail page constitutes trademark infringement.
[0084] Through this confidence-based judgment mechanism, high-confidence infringement and possible misjudgment situations can be effectively distinguished, thereby protecting the legitimate rights and interests of trademark right holders while avoiding unnecessary distress to merchants.
[0085] Step S5260, when the detection result represents that the preset trademark brand logo exists in the product picture and the confidence is less than the confidence threshold, it is determined whether the text information in the product picture contains the trademark text information, and if so, it is determined that the product detail page is trademark infringement.
[0086] When the logo detection model analyzes the product picture, the output detection result shows that the preset trademark brand logo exists in the picture, but the corresponding confidence is lower than the preset confidence threshold, at this time, further verification process is entered. At this time, it is due to the fact that the logo in the picture is not clear enough, there are interference factors, or the form of the logo is different from the training sample, which leads to low confidence of the model on the output result. By further checking the text information in the product picture, the accuracy of the judgment is improved.
[0087] In the product picture, text information exists in various forms, such as watermark text, label text, and advertising language in the picture. These text information contains trademark brand names or other trademark-related content. In some embodiments, the text information in the picture is extracted by optical character recognition (OCR) technology, and compared with the preset trademark text information. The OCR technology can convert the text in the picture into an editable text format, thereby facilitating subsequent text analysis. Assuming that there is a not very clear trademark logo in the product picture, the confidence of the logo detection model is 0.85, which is lower than the preset confidence threshold 0.90. At this time, the OCR technology is started to extract the text in the picture. If the OCR technology identifies that the picture contains the trademark brand name "Brand A", the similarity between the extracted text and the preset trademark text information is further calculated. For example, by calculating the edit distance or cosine similarity, it can be determined whether the text content is highly matched with the trademark text. If the similarity between the text information extracted by the OCR and the trademark text information exceeds the preset text similarity threshold (for example, 0.8), it is determined that the text information in the product picture contains the trademark text information, thereby determining that the product detail page constitutes trademark infringement. This double verification mechanism combining picture logo detection and text information analysis can effectively make up for the shortcomings of a single detection method, and improve the accuracy and reliability of the infringement detection.
[0088] In this embodiment, the problem of infringement detection under low confidence can be effectively handled, avoiding misjudgment or missed judgment caused by unclear logo or interference factors. By analyzing the text information in the picture through the OCR technology, the existence of the trademark is further confirmed, thereby protecting the legitimate rights and interests of the trademark right holder while reducing the misjudgment of the merchant. This double verification method combining visual recognition and text analysis not only improves the accuracy of detection, but also enhances the robustness of the infringement product detection method of the e-commerce platform, making it adaptable to more complex scenarios.
[0089] On the basis of any embodiment of the method of the present application, before the product picture is input into the logo detection model to obtain the detection result of whether the preset trademark brand logo exists in the product picture and the corresponding confidence, the following steps are included:
[0090] Step S6100, a sample image set containing a preset trademark brand logo is obtained, and the sample image set is labeled with a trademark brand logo.
[0091] A large number of sample images containing pre-set trademark brand logos are collected to train a detection model capable of accurately identifying trademark brand logos. These sample images must contain various pre-set trademark brand logos to ensure that the model can learn the characteristics of different trademark logos. The sources of sample images can be diverse, including but not limited to brand official pictures, authorized advertising pictures, and common commodity packaging pictures on the market, etc. These pictures need to cover different scenes, angles, lighting conditions and backgrounds to enhance the generalization ability of the model.
[0092] After collecting the sample images, further detailed labeling is performed on these sample images. The purpose of labeling is to clearly indicate the specific location, shape, size and brand name of the trademark brand logo in each image. For example, if the sample image contains the trademark logo of "Brand A", the labeler needs to label the bounding box of the logo in the image and mark the brand name as "Brand A". This labeling process usually needs professional labeling tools to complete to ensure the accuracy and consistency of labeling. It should be noted that detailed labeling specifications need to be developed during the labeling process to avoid inconsistencies in labeling due to subjective differences among labelers. For example, the labeling specification specifies the drawing method of the bounding box and the labeling format of the brand name, etc.
[0093] In addition to labeling the location and name of the trademark brand logo, other features of the trademark logo such as color, font, deformation degree, etc. can be further labeled. These additional labeling information can help the logo detection model to better learn the visual features of the trademark logo, thereby improving the detection accuracy.
[0094] Step S6200, based on the labeled sample image set, train the logo detection model to a converged state, so that the logo detection model learns the ability to detect whether there is a trademark brand logo in the input commodity picture.
[0095] After collecting and annotating a set of sample images containing trademark brand identifiers in the previous step, these annotated images are used to train an identifier detection model. The identifier detection model is built based on deep learning techniques, such as Convolutional Neural Networks (CNN), which have shown great performance in image recognition tasks and can automatically extract features from images. At the beginning of the training process, the identifier detection model randomly initializes its parameters. Then, the identifier detection model receives annotated sample images batch by batch and tries to predict the location and class of trademark brand identifiers in each image. Since these sample images have been annotated, the identifier detection model can evaluate its performance by comparing its own predictions with the annotations. This comparison is achieved through a loss function, which calculates the difference between the predicted results and the true annotations and outputs a numerical value representing the size of this difference. For example, if the location of the trademark identifier predicted by the identifier detection model is significantly different from the actual annotated location, the value of the loss function will be higher.
[0096] After each prediction, the identifier detection model adjusts its parameters based on the value of the loss function to reduce the prediction error, i.e., backpropagation. By continuously adjusting the parameters, the identifier detection model gradually learns how to more accurately identify trademark brand identifiers. During the training process, the identifier detection model goes through multiple iteration cycles, each of which includes multiple passes through the entire set of sample images. As training progresses, the parameters of the identifier detection model are gradually optimized, and the value of the loss function gradually decreases, indicating that the prediction results of the identifier detection model are increasingly close to the true annotation information.
[0097] During the training process, some techniques are also used to prevent the model from overfitting. Overfitting refers to the situation where the model performs well on the training set but performs poorly on new, unseen data. To avoid overfitting, data augmentation techniques are used to generate more training data by rotating, scaling, cropping, and adjusting the color of training samples, thereby enhancing the generalization ability of the model. In addition, regularization techniques such as L2 regularization are used to limit the complexity of model parameters by adding a regularization term to the loss function, thereby reducing the risk of overfitting.
[0098] When the training of the identifier detection model reaches a certain level, it will enter a state of convergence. Convergence state refers to the situation where the parameters of the model change very little in consecutive iterations, or the value of the loss function no longer decreases significantly. At this point, the identifier detection model has learned the ability to detect whether there is a trademark brand identifier in the input image of the goods.
[0099] In this embodiment, the accuracy and generalization ability of trademark brand identification detection can be effectively improved. Through comprehensive sample image labeling and deep learning model training, the model not only can identify the location and category of the trademark identification, but also can learn various visual features of the trademark identification, so as to accurately identify the trademark identification in different scenes, angles, lighting conditions and backgrounds. In addition, through data enhancement and regularization technology, the generalization ability of the model is significantly improved, so that it can better adapt to unseen data, thereby showing higher robustness in actual application.
[0100] On the basis of any embodiment of the method of the present application, it is judged whether the commodity picture contains a preset trademark brand identification. When it contains, it is determined that the commodity detail page constitutes trademark infringement, comprising:
[0101] Step S7100, image segmentation is performed on the commodity picture according to a preset size specification, and a plurality of identification graph elements are obtained;
[0102] Before image segmentation, the granularity of segmentation is determined according to the preset size specification. These size specifications are set based on the common size of the trademark brand identification and the actual resolution of the picture, so as to ensure that the segmented identification graph elements can contain sufficient detailed information for subsequent detection. For example, if the trademark identification is small, a smaller segmentation size is selected, and if the trademark identification is large, a larger segmentation size is selected.
[0103] After obtaining the commodity picture, the picture is uniformly segmented according to the preset size specification. This process is realized by image processing algorithm, for example, using sliding window technology. The sliding window will move pixel by pixel on the commodity picture according to the set size, and each movement will intercept an image area with the same size as the window, thereby generating a plurality of identification graph elements.
[0104] The segmented identification graph elements will be extracted one by one for subsequent trademark brand identification detection. These identification graph elements contain various parts of the commodity, such as the product itself, the packaging, the background, etc. By segmenting the picture into multiple graph elements, each area in the picture can be analyzed in more detail, thereby avoiding the situation that the trademark identification is missed due to its small size or unclear location.
[0105] In some embodiments, some optimization algorithms are used to improve the efficiency of image segmentation. For example, if the background of the commodity picture is relatively simple, the edge detection algorithm is used to locate the contour of the commodity first, and then the segmentation is performed on the commodity area, thereby reducing unnecessary calculation. This method not only can improve the segmentation speed, but also can improve the detection accuracy, because the background area usually does not contain trademark identification.
[0106] Step S7200, feature matching of the plurality of logo elements with the trademark brand logo in the logo database, screening out logo elements with similarity exceeding a preset fourth similarity threshold value;
[0107] After the image segmentation in the previous step, a plurality of logo elements are obtained, each of which is a local region of the product picture. These logo elements need to be feature matched with the trademark brand logo in the logo database to determine whether the trademark brand logo is contained in these logo elements. The logo database is an image set containing registered trademark brand logos, which are carefully processed and labeled and used as reference standards for detection.
[0108] Before feature matching, image features are extracted from each logo element. The image features include color distribution, texture features, shape features, etc. For example, a color histogram can be used to describe the color distribution of the element, a texture analysis algorithm can be used to extract texture features, or an edge detection algorithm can be used to extract shape features, which can describe the content of the image from different angles and provide rich information for subsequent matching. Meanwhile, the same features are extracted from the trademark brand logo in the logo database, so that each trademark brand logo also has a corresponding set of feature vectors.
[0109] The features of the logo elements are compared with the features of the trademark brand logo in the logo database one by one. The comparison is realized by calculating the similarity between the feature vectors. For example, the Euclidean distance, Hamming distance or cosine similarity algorithm is used to measure the similarity between two feature vectors, and the logo elements with similarity exceeding a preset fourth similarity threshold value are screened out. The preset fourth similarity threshold value is set according to the actual application scenario and the requirement for detection accuracy. A higher threshold value can reduce false positives, but may miss some infringing logos with lower similarity; a lower threshold value can improve the sensitivity of detection, but may increase the false positive rate, therefore, the similarity threshold value needs to be reasonably set according to the characteristics of the trademark logo and actual demand.
[0110] Step S7300, determining the image center distance between each of the screened logo elements in the product picture, and calculating the average center distance;
[0111] Further confirmation of whether the logo elements screened in the previous step constitute infringement behavior requires analysis of the spatial position relationship of these logo elements in the picture. Specifically, the image center position of each logo element is determined. The image center position is obtained by simple geometric calculation and is the center point of the logo element bounding box. For example, if the position of a logo element in the picture is bounded by a rectangular box, the center point of the rectangular box is taken as the image center of the element. In this way, a clear center position can be determined for each screened logo element.
[0112] Then the image center distance between each two identified image elements is calculated. The calculation of the image center distance is based on the coordinates of the image center, i.e. the straight-line distance between each pair of image element center points is measured to determine the distance between all the selected identified image elements.
[0113] After calculating the image center distance between all the pairs, the average value of the image center distance is calculated by adding all the distances and dividing by the number of distances. By calculating the average value of the center distance, the distribution of the selected identified image elements in the product picture can be evaluated. If the average distance is small, it indicates that the identified image elements are more concentrated in the picture, and if the average distance is large, it indicates that the identified image elements are more dispersed in the picture.
[0114] Step S7400, when the average value of the center distance is less than a preset distance threshold, it is determined that the product detail page constitutes trademark infringement.
[0115] This step compares the average value of the center distance with a preset distance threshold. The preset distance threshold is set by a person skilled in the art according to the typical layout of the trademark logo and the actual application scenario, and is used to distinguish between normal trademark use and possible infringement. When the similarity of these identified image elements to the trademark brand logo has exceeded the preset fourth similarity threshold, and the average value of the center distance of these identified image elements in the product picture is less than the preset distance threshold, it indicates that these identified image elements are more concentrated. When a user combines two authorized trademark logos to evade infringement detection, this step can be used to identify and determine this potential infringement, i.e. because the user intentionally combines or juxtaposes multiple trademark logos to achieve the purpose of confusing the audience and evading infringement detection. For example, the user places two similar trademark logos very close together, or places the parts of the trademark logo in a small spatial range while keeping them apart, trying to evade single logo infringement detection in this way. The determination mechanism of this step can effectively identify the situation of evading detection by abnormal combination or layout of trademark logos, thereby ensuring the comprehensiveness and accuracy of trademark infringement detection.
[0116] In this embodiment, the trademark infringement problem in complex scenarios can be effectively handled, especially when the infringer tries to evade detection by combining or juxtaposing multiple trademark logos. Through the combination of image segmentation and spatial distribution analysis, the present application can not only identify single logo infringement, but also detect abnormal combinations of multiple logos.
[0117] On the basis of any embodiment of the method of the present application, the relevance of the commodity description text to the preset copyright description text set and the relevance of the commodity picture to the preset copyright picture set are determined, and when any of the relevances meets a preset condition, it is determined that the commodity detail page constitutes copyright infringement, including:
[0118] In step S5310, based on the commodity description text, the text feature vector with the highest similarity is retrieved from the first vector database containing the copyright description text set, and the corresponding fourth similarity is determined.
[0119] In an embodiment, an efficient vector retrieval system is constructed for storing and retrieving feature vectors of copyright description texts, and the core of the vector retrieval system is the first vector database, which can quickly process large-scale vector data and support efficient similarity retrieval. In this step, FAISS (Facebook AI Similarity Search) can be used as the vector retrieval system. FAISS is an open source library designed for large-scale vector retrieval and clustering, supporting multiple distance measurement methods such as cosine similarity and Euclidean distance, and can quickly find the most similar vector to the target vector in a large amount of data, i.e. it can quickly retrieve the text feature vector with the highest similarity in the first vector database. In order to ensure the accuracy of the retrieval, the copyright description text needs to be preprocessed and feature extracted. Preprocessing includes steps such as text cleaning and word segmentation. For example, the jieba word segmentation tool can be used to segment the copyright description text, and after preprocessing, the copyright description text needs to be feature extracted to convert the text into a numerical vector that can describe its semantics. In this embodiment, the BERT (Bidirectional Encoder Representations from Transformers) model is used for feature extraction. BERT is a pre-trained language model based on the Transformer architecture, which can capture bidirectional semantic information in the text. By inputting the copyright description text into the BERT model, the feature vector of each text can be obtained, which can reflect the semantic content of the text, such as theme, sentiment, style, etc.
[0120] In some embodiments, the BERT model is fine-tuned to further improve the quality of the text feature vector, where fine-tuning refers to further training the BERT model using data in a specific field based on the pre-trained model to make it better adapt to the current task. In this step, the copyright description text is used as fine-tuning data to train the BERT model, so that it can better capture the semantic features of the copyright text.
[0121] For the product description text of the target product, preprocessing and feature extraction are also needed, and finally the corresponding target feature vector is obtained. Then the target feature vector obtained is subjected to similarity retrieval with the copyright description text feature vector in the first vector database. The vector retrieval system can quickly calculate the similarity between the target feature vector and all text feature vectors in the database, and return the most similar text feature vector and its similarity value, i.e. the fourth similarity value. In some embodiments, cosine similarity is used as the similarity measurement method. The value of cosine similarity ranges between [-1, 1], and the closer the value is to 1, the more similar the two vectors are.
[0122] Through this step, the semantic similarity between the product description text and the copyright description text can be used to accurately determine whether the product detail page constitutes copyright infringement. This not only allows for fast processing of large-scale text data, but also effectively captures the semantic information of the text, providing an efficient and accurate technical means for detecting copyright infringement of e-commerce platform product detail pages.
[0123] Step S5320, based on the product picture, retrieve the image feature vector with the highest similarity from the second vector database containing the copyright picture set, and determine the corresponding fifth similarity;
[0124] In one embodiment, a vector retrieval system for storing and retrieving copyright picture feature vectors is constructed. The core of the vector retrieval system is the second vector database. The constructed second vector database can process large-scale image feature data and support efficient similarity retrieval. In this embodiment, FAISS (Facebook AI Similarity Search) is used as the vector retrieval system. Before constructing the vector retrieval system, the copyright picture set needs to be preprocessed and feature extracted. Preprocessing includes adjusting the resolution, cropping, and normalizing the pictures, etc. to ensure that all pictures have consistent format and size. For example, all copyright pictures are uniformly adjusted to 224x224 pixels in size so that the subsequent feature extraction model can process them. At the same time, the pictures are normalized, i.e. the pixel value range is adjusted to [0, 1] or [-1, 1] to reduce the impact of the value range on feature extraction.
[0125] After the preprocessing is completed, feature extraction is performed on the copyright pictures to convert the pictures into numerical vectors that can describe their visual content. In one embodiment, a pre-trained convolutional neural network (CNN) model is used for feature extraction. For example, models such as ResNet, VGG, or CLIP can be used. These models have been pre-trained on large-scale image datasets and can extract image features with semantic meaning. Taking the CLIP model as an example, the CLIP model is a multi-modal model that can process both image and text data and extract semantic associations between images and text. When extracting image features, the copyright pictures are input into the CLIP model, which outputs a fixed-length feature vector that can capture the visual content and semantic information of the pictures, such as the shapes, colors, textures, and scenes of the objects in the pictures.
[0126] For the pictures of the target product, preprocessing and feature extraction are also required to obtain its feature vector. Then, similarity retrieval is performed between the feature vector and the feature vectors of the copyright pictures in the second vector database. FAISS can quickly calculate the similarity between the target vector and all vectors in the database and return the most similar vector and its similarity value.
[0127] Step S5330, when the fourth similarity is greater than the preset text similarity threshold or the fifth similarity is greater than the preset image similarity threshold, it is determined that the product detail page is copyright infringement.
[0128] The text similarity threshold and the image similarity threshold are pre-set and are key parameters for determining whether the product detail page constitutes infringement. They are set by those skilled in the art according to the actual application scenario and the needs of copyright protection. During the infringement detection process, both the text similarity and the image similarity need to be considered. Specifically, when the fourth similarity between the product description text and the copyright description text is greater than the preset text similarity threshold, or the fifth similarity between the product picture and the copyright picture is greater than the preset image similarity threshold, it can be determined that the product detail page constitutes copyright infringement. This comprehensive evaluation method can more comprehensively cover various forms of infringement behavior and avoid misjudgment or omission in a single dimension.
[0129] In this embodiment, copyright infringement detection covering both text and image dimensions is achieved. By combining the semantic similarity of the text and the visual similarity of the image, an efficient and accurate infringement product detection method is provided. This embodiment not only can quickly process large-scale text and image data, but also can effectively capture the semantic information of the text and image, thereby showing higher accuracy and reliability in protecting copyrights.
[0130] Before retrieving the text feature vector with the highest similarity from the first vector database containing a set of copyright description texts and determining the corresponding fourth similarity based on the commodity description text in any embodiment of the method of the application, the following steps are included:
[0131] In step S8100, a set of preset copyright description texts is obtained, each copyright description text in the set of copyright description texts is subjected to word segmentation processing and semantic vectorization processing, a corresponding text feature vector is generated, and the copyright description text is stored in the first vector database in association with the text feature vector.
[0132] In an embodiment, a corpus containing copyright description texts is constructed, which covers various copyright text contents related to the target commodity. These copyright text contents can be obtained from documents provided by copyright parties, official websites, authorized e-commerce platform pages, etc. The obtained text data needs to be cleaned and preprocessed to remove irrelevant symbols, stop words and repetitive content, so as to ensure the quality and consistency of the corpus. Then, the copyright description texts are subjected to word segmentation processing. After word segmentation, each segmented text unit is subjected to semantic vectorization processing. Pre-trained language models based on deep learning, such as BERT and RoBERTa, are used. These models can capture the semantic information of the text through unsupervised learning on large-scale corpus. The specific operation is to input the segmented text into the pre-trained model to obtain the context-related embedding vector of each word, and then combine the vectors of multiple words into the feature vector of the entire text through aggregation operations such as taking the average value, maximum pooling, etc. In addition, traditional methods such as TF-IDF can be combined to weight the vectors to highlight the contribution of important words.
[0133] The generated text feature vector needs to be associated with the corresponding copyright description text and stored in the first vector database. The database uses efficient vector retrieval systems such as FAISS (Facebook AI Similarity Search) or Milvus, etc. These systems can quickly perform similarity retrieval on large-scale vector data and support multiple distance measurement methods such as cosine similarity and Euclidean distance. During storage, a unique identifier needs to be set for each text feature vector, and the source, category and other meta information of the corresponding copyright description text need to be recorded, so that it can be quickly located and traced back during subsequent retrieval.
[0134] Through this step, a high-quality and efficient copyright description text vector database is constructed, which provides a solid foundation for subsequent infringement detection based on commodity description texts. In the infringement detection process, when it is necessary to determine whether the commodity description text constitutes copyright infringement, the text feature vector with the highest similarity to the commodity description text can be directly retrieved from the vector database, and whether there is an infringement can be determined according to the similarity threshold.
[0135] In step S8200, a preset copyright picture set is obtained, image feature extraction processing is performed on each copyright picture in the copyright picture set, a corresponding image feature vector is generated, and the corresponding copyright picture is stored in the second vector database.
[0136] The copyright picture set contains all copyright pictures related to the target product, which come from the official picture library provided by the copyright party, the authorized e-commerce platform page, the brand promotion material, etc. The collected copyright pictures should cover different angles, backgrounds and lighting conditions to ensure the comprehensiveness and robustness of feature extraction. After collecting the copyright picture set, each picture needs to be preprocessed, and then the preprocessed copyright picture is subjected to image feature extraction. In this step, a deep learning model can be used to extract image features. For example, a pre-trained convolutional neural network (CNN) model such as ResNet, VGG or CLIP can be used. These models have been pre-trained on large-scale image datasets and can extract image features with semantic meaning.
[0137] The generated image feature vector needs to be associated with the corresponding copyright picture and stored in the second vector database. This database can use an efficient vector retrieval system such as FAISS or Milvus, which can quickly perform similarity retrieval in large-scale vector data and support multiple distance measurement methods. In the storage process, a unique identifier needs to be set for each image feature vector, and the source, category and other meta information of the corresponding copyright picture need to be recorded, so that it can be quickly located and traced back during subsequent retrieval.
[0138] Through this step, a high-quality and efficient copyright picture feature vector database is constructed. When it is necessary to determine whether the product picture constitutes copyright infringement, the image feature vector with the highest similarity to the product picture can be directly retrieved from the vector database, and whether there is an infringement can be determined according to the similarity threshold.
[0139] In this embodiment, both text and picture copyright content can be processed simultaneously, and a high-quality and efficient copyright feature vector database is constructed. Through the efficient vector retrieval system (such as FAISS or Milvus), fast similarity retrieval is realized, and multiple distance measurement methods are supported, so that it can quickly determine whether the product description text and product picture constitute copyright infringement.
[0140] Please refer to Figure 3According to one aspect of the present application, a device for detecting infringing goods is provided, comprising an event response module 5100, a trademark infringement detection module 5200, a copyright infringement detection module 5300, and a result processing module 5400. The event response module 5100 is configured to respond to a target product listing event triggered by a user, start a pre-infringement detection process in a listing process of the target product, and obtain product description text and product pictures in a product detail page of the target product. The trademark infringement detection module 5200 is configured to determine whether the product description text contains segmented words with a similarity greater than a preset similarity threshold with respect to a preset trademark text information, or whether the product pictures contain a preset trademark brand logo, and determine that the product detail page constitutes trademark infringement when the product description text contains the segmented words or the product pictures contain the trademark brand logo. The copyright infringement detection module 5300 is configured to determine the relevance between the product description text and product pictures and a preset copyright description text set and copyright picture set, and determine that the product detail page constitutes copyright infringement when any of the relevances satisfies a preset condition. The result processing module 5400 is configured to construct an infringing detection report corresponding to the product detail page when the product detail page is determined to constitute trademark infringement or copyright infringement, and push the infringing detection report to the user and end the listing process of the target product.
[0141] On the basis of any embodiment of the device of the present application, the trademark infringement detection module 5200 comprises a text segmentation sub-module configured to perform segmentation processing on the product description text to obtain a plurality of target segmented words, the product description text including a product title, product detail information, and brand information; a similarity calculation sub-module configured to calculate a first similarity between the target segmented words corresponding to the product title and the trademark text information, a second similarity between the target segmented words corresponding to the product detail information and the trademark text information, and a third similarity between the target segmented words corresponding to the brand information and the trademark text information; and a trademark infringement determination sub-module configured to determine that the target product detail page contains trademark infringement when the first similarity is greater than a preset first similarity threshold, the second similarity is greater than a preset second similarity threshold, or the third similarity is greater than a preset third similarity threshold.
[0142] On the basis of any embodiment of the device of the present application, the trademark infringement detection module 5200 further comprises: an image mark detection submodule, configured to input the product picture into a mark detection model to obtain a detection result of whether a preset trademark brand mark exists in the product picture and a corresponding confidence degree; a high-confidence degree determination submodule, configured to determine that the product detail page constitutes trademark infringement when the detection result indicates that the preset trademark brand mark exists in the product picture and the confidence degree is greater than or equal to a preset confidence degree threshold; and a low-confidence degree verification submodule, configured to determine whether text information in the product picture contains the trademark text information when the detection result indicates that the preset trademark brand mark exists in the product picture and the confidence degree is less than the confidence degree threshold, and determine that the product detail page constitutes trademark infringement if the text information contains the trademark text information.
[0143] On the basis of any embodiment of the device of the present application, the image mark detection submodule comprises: an image set labeling submodule, configured to obtain a sample image set containing a preset trademark brand mark and label the sample image set for the trademark brand mark; and a model training submodule, configured to train the mark detection model to a convergent state based on the labeled sample image set, so that the mark detection model learns the ability to detect whether a trademark brand mark exists in an input product picture.
[0144] On the basis of any embodiment of the device of the present application, the trademark infringement detection module 5200 further comprises: an image segmentation processing submodule, configured to perform image segmentation on the product picture according to a preset size specification to obtain a plurality of mark graph elements; a trademark feature matching submodule, configured to perform feature matching between the plurality of mark graph elements and trademark brand marks in a mark database to screen out mark graph elements with a similarity exceeding a preset fourth similarity threshold; a spatial distribution analysis submodule, configured to determine image center distances between each two of the screened mark graph elements in the product picture and calculate an average center distance; and a spatial distribution determination submodule, configured to determine that the product detail page constitutes trademark infringement when the average center distance is less than a preset distance threshold.
[0145] On the basis of any embodiment of the device of the present application, the copyright infringement detection module 5300 comprises: a text feature retrieval submodule configured to retrieve the highest-similarity text feature vector from a first vector database containing a set of copyright description texts based on the commodity description text, and determine a corresponding fourth similarity; an image feature retrieval submodule configured to retrieve the highest-similarity image feature vector from a second vector database containing a set of copyright pictures based on the commodity picture, and determine a corresponding fifth similarity; and a copyright infringement judgment submodule configured to determine that the commodity detail page is copyright infringement when the fourth similarity is greater than a preset text similarity threshold or the fifth similarity is greater than a preset image similarity threshold.
[0146] On the basis of any embodiment of the device of the present application, the text feature retrieval submodule comprises: a text feature library construction submodule configured to obtain a set of preset copyright description texts, perform word segmentation processing and semantic vectorization processing on each copyright description text in the set of copyright description texts, generate a corresponding text feature vector, and store the copyright description text associated with the text feature vector in the first vector database; and an image feature library construction submodule configured to obtain a set of preset copyright pictures, perform image feature extraction processing on each copyright picture in the set of copyright pictures, generate a corresponding image feature vector, and store the copyright picture associated with the image feature vector in the second vector database.
[0147] To solve the above technical problems, the embodiments of the present application further provide a computer device. As shown in the internal structure diagram of the computer device. Figure 4 The computer device comprises a processor, a computer readable storage medium, a memory and a network interface connected through a system bus. The computer readable storage medium of the computer device stores an operating system, a database and computer readable instructions. The database can store a control information sequence. When the computer readable instructions are executed by the processor, the processor can implement an infringement commodity detection method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer readable instructions. When the computer readable instructions are executed by the processor, the processor can execute the infringement commodity detection method of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that the structure shown in the Figure 4 application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0148] In the present embodiment, the processor is configured to execute Figure 3The specific functions of each module and its sub-modules in the above embodiment, the memory stores the program code and various data required to execute the above-mentioned modules or sub-modules. The network interface is used for data transmission between the user terminal or the server. The memory in the above embodiment stores the program code and data required to execute all modules / sub-modules of the infringement product detection device of the present application, and the server can call the program code and data of the server to execute the functions of all sub-modules.
[0149] The present application also provides a storage medium having computer readable instructions stored therein, which, when executed by one or more processors, cause the one or more processors to perform the steps of the infringement product detection method of any embodiment of the present application.
[0150] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments of the present application can be completed by a computer program instructing related hardware. The computer program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of each method. The storage medium can be a computer readable storage medium such as a magnetic disc, an optical disc, a read-only memory (ROM), or a random access memory (RAM).
[0151] A person of ordinary skill in the art can understand that the steps, measures, and schemes in the various operations, methods, and processes discussed in the present application can be alternated, changed, combined, or deleted. Further, other steps, measures, and schemes in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and schemes in the prior art with the various operations, methods, and processes disclosed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0152] The above only describes some embodiments of the present application. It should be noted that for a person of ordinary skill in the art, without departing from the principles of the present application, several improvements and refinements can be made, which should also be considered as the protection scope of the present application.
Claims
1. A method for detecting infringing goods, characterized in that: include: In response to a user-triggered target product listing event, the pre-infringement detection process in the target product listing process is initiated to obtain the product description text and product images on the target product's product details page; Determine whether the product description text contains a word whose similarity with the preset trademark text information is greater than a preset similarity threshold, or determine whether the product image contains a preset trademark brand logo, and when it is contained, determine that the product details page constitutes trademark infringement, wherein determining whether the product image contains a preset trademark brand logo, and when it is contained, determine that the product details page constitutes trademark infringement, including: performing image segmentation on the product image according to preset size specifications to obtain multiple identification primitives; performing feature matching on the multiple identification primitives with the trademark brand logos in the identification database, and screening out identification primitives whose similarity exceeds a preset fourth similarity threshold; determining the image center distance between each of the screened identification primitives in the product image, and calculating the average center distance; when the average center distance is less than the preset distance threshold, determine that the product details page constitutes trademark infringement; Determining the relevance of the product description text to a preset set of copyright description texts, and the relevance of the product image to a preset set of copyright images, and determining that the product details page constitutes copyright infringement when either relevance meets a preset condition; When the product details page is judged to be a trademark infringement or copyright infringement, a corresponding infringement detection report is constructed and pushed to the user, ending the listing process of the target product.
2. The method for detecting infringing goods according to claim 1, characterized in that: Determining whether the product description text contains a word whose similarity to the preset trademark text information is greater than a preset similarity threshold, and if so, determining that the product details page constitutes trademark infringement, including: Performing word segmentation processing on the product description text to obtain multiple target word segments, wherein the product description text includes a product title, product details information, and brand information; Calculating respectively a first similarity between the target segmentation word corresponding to the product title and the trademark text information, a second similarity between the target segmentation word corresponding to the product details information and the trademark text information, and a third similarity between the target segmentation word corresponding to the brand information and the trademark text information; When the first similarity is greater than a preset first similarity threshold, the second similarity is greater than a preset second similarity threshold, or the third similarity is greater than a preset third similarity threshold, it is determined that there is trademark infringement on the target product details page.
3. The method for detecting infringing goods according to claim 1, characterized in that: Determine whether the product image contains a preset trademark brand logo. If so, determine that the product details page constitutes trademark infringement, or achieve this through the following steps: Input the product image into the logo detection model to obtain a detection result of whether a preset trademark brand logo exists in the product image and the corresponding confidence level; When the detection result indicates that a preset trademark brand logo exists in the product image and the confidence level is greater than or equal to a preset confidence level threshold, the product details page is determined to be trademark infringing; When the detection result indicates that a preset trademark brand logo exists in the product image and the confidence level is less than the confidence level threshold, determine whether the text information in the product image contains the trademark text information. If so, determine that the product details page is a trademark infringement.
4. The method for detecting infringing goods according to claim 3, characterized in that: Before inputting the product image into the logo detection model to obtain a detection result of whether a preset trademark brand logo exists in the product image and a corresponding confidence level, the method includes: Acquire a sample image set containing a preset trademark brand logo, and label the sample image set with the trademark brand logo; Based on the labeled sample image set, the logo detection model is trained to convergence, so that the logo detection model learns the ability to detect whether there is a trademark brand logo in the input product image.
5. The method for detecting infringing goods according to claim 1, characterized in that: Determining the relevance of the product description text to a preset set of copyright description texts, and the relevance of the product image to a preset set of copyright images, and determining that the product details page constitutes copyright infringement when either relevance meets a preset condition, including: Based on the product description text, retrieve the text feature vector with the highest similarity from a first vector database containing a set of copyright description texts, and determine the corresponding fourth similarity; Based on the product image, retrieving the image feature vector with the highest similarity from a second vector database containing a set of copyrighted images, and determining the corresponding fifth similarity; When the fourth similarity is greater than a preset text similarity threshold or the fifth similarity is greater than a preset image similarity threshold, it is determined that the product details page is copyright infringing.
6. The method for detecting infringing goods according to claim 5, characterized in that: Based on the product description text, before retrieving the text feature vector with the highest similarity from a first vector database containing a set of copyright description texts and determining the corresponding fourth similarity, the method includes: Obtaining a preset copyright description text set, performing word segmentation and semantic vectorization on each copyright description text in the copyright description text set, generating a corresponding text feature vector, associating the copyright description text, and storing the vector in the first vector database; A preset copyright image set is obtained, image feature extraction processing is performed on each copyright image in the copyright image set, a corresponding image feature vector is generated, the corresponding copyright image is associated, and the vector is stored in the second vector database.
7. A device for detecting infringing goods, characterized in that: include: An event response module is configured to respond to a user-triggered target product listing event, initiate a pre-infringement detection process in the target product listing process, and obtain product description text and product images from the target product's product details page; The trademark infringement detection module is configured to determine whether the product description text contains a word whose similarity with the preset trademark text information is greater than a preset similarity threshold, or to determine whether the product image contains a preset trademark brand logo. If so, the product details page is determined to constitute trademark infringement. The module comprises: performing image segmentation on the product image according to preset size specifications to obtain multiple identification primitives; performing feature matching on the multiple identification primitives with the trademark brand logos in the identification database to screen out identification primitives whose similarity exceeds a preset fourth similarity threshold; determining the image center distance between each of the screened identification primitives in the product image, and calculating the average center distance; and determining that the product details page constitutes trademark infringement when the average center distance is less than the preset distance threshold. a copyright infringement detection module configured to determine the relevance of the product description text to a preset set of copyright description texts and the relevance of the product image to a preset set of copyright images, and to determine that the product details page constitutes copyright infringement when either relevance satisfies a preset condition; The result processing module is configured to construct a corresponding infringement detection report and push it to the user when the product details page is judged to be a trademark infringement or copyright infringement, thereby ending the listing process of the target product.
8. An infringing goods detection device, comprising a central processing unit and a memory, characterized in that: The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 6.
9. A non-volatile readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 6 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.
Citation Information
Patent Citations
Identification method, computing device, computer storage medium and computer program product
CN117910450A