A system of trademark surveillance system with image understanding for multi-country and the method thereof

The multilingual trademark monitoring system addresses inefficiencies in existing systems by employing image semantic understanding and machine learning to automate trademark monitoring, enhancing accuracy and reducing costs through automated, customizable reporting.

WO2026019427A1PCT designated stage Publication Date: 2026-01-22WU PENG CHUN +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/038615
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing trademark monitoring systems face inefficiencies and inaccuracies due to manual division of graphic element codes, subjective judgments, and limited ability to handle graphic trademarks, leading to inconsistent results and increased operational costs.

Method used

A multilingual trademark monitoring system utilizing image semantic understanding, including an input module, information storage, information understanding module, monitoring comparison module, and report generation module, which performs text and image analysis to compare and generate monitoring reports based on machine learning models trained with trademark data.

Benefits of technology

Enhances accuracy and efficiency in trademark monitoring by automating the process, providing comprehensive and customizable reports, and improving the handling of graphic trademarks, reducing human error and operational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024038615_22012026_PF_FP_ABST
    Figure US2024038615_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a system of trademark surveillance system with image understanding for multi-country and the method thereof, which is implemented by providing a user-operated user electronic device. The user electronic device comprises a second processor and a second network interface controller. A server comprises a second application program. The second processor is connected to the server through the second network interface controller and executes the second application program to perform trademark monitoring, particularly to monitor trademark images by inputting a text description.
Need to check novelty before this filing date? Find Prior Art

Description

A SYSTEM OF TRADEMARK SURVEILLANCE SYSTEM WITH IMAGE UNDERSTANDING FOR MULTI-COUNTRY AND THE METHOD THEREOFTECHNICAL FIELD

[0001] The present invention relates to a system of trademark surveillance system with image understanding for multi-country and the method thereof, particularly by monitoring trademarks through text understanding, image understanding, and image-to-text description. Furthermore, it can also monitor images corresponding to a text description.BACKGROUND OF THE INVENTION

[0002] Trademark search is crucial for trademark registration applications, trademark review, trademark management and trademark rights protection. Traditional graphic trademark retrieval essentially achieves the search purpose by manually inputting trademark graphic element codes as search conditions. The trademark graphical element code is a classification tool for trademark graphical elements generated based on the Vienna Agreement on the Establishment of an International Classification of Elements of Marks for Goods and Services. It consists of a list of trademark graphical elements classified by main categories, subcategories and groups, including trademark Graphic element number and name. Therefore, each trademark graphic element code represents the content and meaning of the trademark graphic element.

[0003] However, currently in national trademark management agencies around the world, the division of trademark graphic element codes is mostly done manually by a small number of professional trademark graphic element code examiners, with basically no assistance from intelligent tools or means. Although the existing manual division method of trademark graphic element codes can complete the task of dividing trademark graphic element codes, there are obvious flaws and deficiencies, which are mainly reflected in the following aspects: 1) The work efficiency of manual division of trademark graphic element codes is low, The workloadis huge. 2) The division of trademark graphic element codes requires strong professionalism. It is difficult for ordinary personnel to fully grasp the method of dividing trademark graphic element codes, which limits the widespread application of graphic trademark retrieval. 3) Even if the trademark graphic element codes are divided by professionals, different professionals will subjectively judge the meaning of the trademark graphic differently, which will lead to inconsistency in the trademark graphic element codes.

[0004] Trademarks are an indispensable factor in the commercial economy. The number of trademark applications every year reaches millions, and trademark information reaches tens of millions. For such a huge group of people, if people judge or review whether two trademarks are similar or similar, When the degree is reached, it all depends on the human eye and subjective consciousness to judge. This has great room for improvement in terms of the operation cycle and the objective stability of the results. Nowadays, it has become a hot issue in this field to study how to use computers and other electronic devices to search and match trademarks. People are constantly trying various computer algorithms and multimedia technologies to achieve automatic retrieval of trademark graphics. In the process of feature extraction and comparison, A lot of experiments and improvements have been made in this link, but there is room for improvement in both the pre-processing and post-result output links.

[0005] Existing trademark monitoring methods mainly rely on users' active inquiries. Users need to regularly query trademark-related information and analyze the found relevant information to determine existing risks or matters that need to be dealt with, so as to achieve trademark monitoring. This method of trademark monitoring is costly and inefficient. Moreover, if users neglect to query, important time limits may be missed, resulting in irreparable losses.

[0006] However, there are already some technologies similar to trademark search and word trademark search.

[0007] As disclosed in Chinese Application No. 202110039372.5, a trademark monitoring method, a trademark monitoring device and an electronic device, the trademark monitoring method includes: determining a monitoring subject based on the user's registration information; determining the first trademark corresponding to the monitoring subject; Monitor the first trademark and send monitoring results to the user. This application solves the problems of high cost, low efficiency and poor reliability of existing trademark monitoring methods by automatically monitoring trademarks and sending monitoring results to users. At the same time, this application can determine the trademarks that need to be monitored based on the user's registration information without the need for the user to manually add them, thereby bringing convenience to the user and improving the user experience.

[0008] However, the above-mentioned Chinese application number 202110039372.5 has several problems that need to be improved. Although the above-mentioned technology mentions the subject matter of trademark monitoring, the content of its instructions is relatively empty, and there is no actual explanation of how to monitor, how to compare, and how to generate monitoring. Results and so on.

[0009] As disclosed in Chinese Application No. 201810292281.0, a trademark graphic retrieval method first obtains accurate primary sorting results based on feature extraction and matching methods that combine multiple methods, and then performs secondary or even multiple retrievals on the initial sorting results. The final ranking is based on the sets generated by secondary or even multiple searches. This method integrates the ranking results of multiple searches of related images. Images that appear more frequently have higher weights, and images that are ranked higher have higher weights, weight, fully exploring the correlation between images, and greatly improving the accuracy of trademark graphic similarity determination; the trademark graphic database in step S101 is set to at least include the same product category as the trademark image to be inspected and / or All similar products have been registered.

[0010] However, there are several problems to be improved in the above-mentioned Chinese application number 201810292281.0. First, the above-mentioned technology belongs to the search and retrieval of trademarks, which is obviously different from the so-called trademark monitoring. Furthermore, the graphic comparison method mentioned in the above-mentioned technology , mainly uses image feature extraction for comparison, which can easily lead to problems of accuracy in feature definition, thus affecting the results of comparison and search.

[0011] As disclosed in Chinese Application No. 201510945834.4, a trademark graphic element identification method, device and system includes: establishing a sample trademark database and establishing a corresponding relationship between the sample trademark and the known graphic element coding division data of the registered graphic trademark. ; Extract and process the image feature information of the sample trademark, and establish the corresponding relationship between the sample trademark and the extracted image feature information; Extract and process the image feature information of the trademark to be identified; Use the image feature information of the trademark to be identified as a search Match the search conditions to find the sample trademark with the highest similarity to the image feature information of the trademark to be identified and the corresponding trademark graphic element code; output the trademark graphic element code corresponding to the sample trademark as the graphic element code of the trademark to be identified. The invention can automatically divide trademark graphics into element codes of trademark graphics.

[0012] However, the above-mentioned Chinese application number 201510945834.4 has several problems that need to be improved. The above-mentioned technology also belongs to the application of trademark search, which is different from the so-called trademark monitoring. The above-mentioned technology uses the positioning of the XY axis to perform so-called feature extraction. The bias is to compare shapes to confirm whether the search results are similar. There is also an accuracy problem in feature extraction, and the comparison of appearance may also produce similar outlines of the outermost graphics, but the actual graphicsare not. Dissimilar results may cause errors or misjudgments.

[0013] For example, the existing trademark monitoring software TradeMerch, please refer to Figures 1 to 6. This platform can only monitor "US trademarks", and users can only enter "text", although a large number of text keywords can be entered depending on the cost. , but we all know that trademarks containing graphics still account for the vast majority, so users can only enter text and the monitoring effect is limited. It can also be found from the picture that this platform cannot accept the input of descriptive text, such as in the picture Enter A CAT ON A ROCK, and the monitoring result is 0. It can only identify words. There is a larger gap in monitoring graphic trademarks, which reduces the accuracy of trademark monitoring. In addition, the list of similar trademarks that pops up based on keyword monitoring is quite confusing, and it is impossible to customize the sorting rules. Logically speaking, the most similar trademarks in the previous case should be listed at the top. For example: the entered word is CATON, but in fact the previous case There are trademarks with the same name, but they are not displayed at the top, not even in the top 10. Furthermore, for example, the trademark of a well-known oil company has a distinctive XX identification. From the perspective of trademark monitoring, The identified XX will be specially monitored to see if there are similar uses. However, in addition to the inability to monitor graphics, the platform cannot monitor even if the text XX is entered (the input character source is shown to be too short), which shows that the platform still has many imperfections.

[0014] From the above description, it can be known that it is necessary to improve or adjust the conventional technology in order to provide a complete, accurate and time-cost-saving trademark monitoring system. In view of this, the inventor of the present invention has tried his best to Research and creation, and finally developed the system and method of the present invention.SUMMARY OF THE INVENTION

[0015] The objective of the present invention is to propose a multilingual trademark monitoring system, particularly one that can monitor trademark images based on a text description, thereby solving the problems existing in the prior art.

[0016] Therefore, in order to achieve the above objectives of the present invention, the present invention provides a multilingual trademark monitoring system based on image semantic understanding, wherein a user operates a user electronic device, and the second processor of the user electronic device is connected to a server through a second network interface controller to execute a second application program for trademark monitoring. The system at least comprises:

[0017] An input module for receiving relevant information for trademark monitoring;

[0018] An information storage unit for storing the relevant information for trademark monitoring received by the input module;

[0019] An information understanding module for receiving the relevant information for trademark monitoring from the information storage unit and analyzing the information;

[0020] A monitoring comparison module for performing comparison analysis in a model database to generate a monitoring comparison result based on the relevant information for trademark monitoring in the information storage unit and the text, image, or text description analyzed by the information understanding module;

[0021] A monitoring database for receiving and storing the monitoring comparison result from the monitoring comparison module; and

[0022] A report generation module for generating a monitoring report based on the relevant information for trademark monitoring in the information storage unit.

[0023] wherein, the information understanding module further comprises a text parsing unit, an image parsing unit, and an image-to-text description unit, and performs natural languageprocessing, image understanding, or image-to-text description on the relevant information for trademark monitoring. The processor connects to the server and executes the application program through the network interface controller, thereby activating the trademark online application module and the classification recommend module. Furthermore, the processor further activates the filing data collect unit, classification item select unit, data transmission unit, word analysis unit, classification item compare unit, and report generation unit.

[0024] The text parsing unit performs natural language processing on the text in the relevant information for trademark monitoring, vectorizes the text, and stores the analysis results in the information storage unit.

[0025] The image parsing unit performs image understanding on the images in the relevant information for trademark monitoring, labels the images with feature tags or image basic information, and stores the analysis results in the information storage unit.

[0026] The image-to-text description unit converts the images in the relevant information for trademark monitoring into descriptive text, textualizes the images, vectorizes the descriptive text, and stores it in the information storage unit.

[0027] The monitoring comparison module performs text similarity comparison, image similarity comparison, or vector similarity comparison in the model database.

[0028] The relevant information for trademark monitoring in the information storage unit includes monitoring standards. The monitoring comparison module compares and analyzes the monitoring comparison results based on the similarity values set by the monitoring standards.

[0029] In this embodiment, a manager operates a manager electronic device, and the first processor of the manager electronic device is connected to a server through a first network interface controller to execute a first application program for data update and model training. The system further comprises:

[0030] A data extraction module for extracting trademark public data from multiple trademark data sources;

[0031] An internal database for storing the trademark public data extracted by the data extraction module; and

[0032] A model training module for receiving the trademark public data stored in the internal database and performing model training, and storing the trained training data in the model database.

[0033] In this embodiment, the data extraction module extracts data from the trademark data sources based on a frequency setting. Preferably, the processor connects to the server through the network interface controller and executes the application program, further configures and activates the risk assessment module to perform prior art comparison on the trademark names in the user-inputted trademark application data. The processor within the application program executes risk assessment module and further configure to activate the text parsing unit, retrieval comparison unit, and report generation unit.

[0034] The data extraction module extracts data using database APIs, web scraping, or filebased extraction.

[0035] The training method of the model training module further includes text training, image training, and image-to-text description training.

[0036] Image-to-text description training involves training the model to learn the relationship between images and text descriptions using image and text description data. The model then uses these relationships to generate text descriptions for new images, enabling it to understand both the semantics of images and text.

[0037] Another objective of the present invention is to propose a multilingual trademark monitoring method based on image semantic understanding to solve the problems existing inthe prior art.

[0038] Therefore, in order to achieve the above objectives of the present invention, the method provided by the present invention is implemented by providing a user-operated user electronic device. The method of monitoring trademarks by operating a user electronic device from the user side to input relevant trademark monitoring information, connecting to a server through a second network interface controller of the user electronic device using a second processor, and executing a second application program, at least comprises the following steps:

[0039] (1) The user inputs and stores relevant monitoring information for trademark monitoring through the input module of the user electronic device into an information storage unit;

[0040] (2) Text analysis, image analysis, or image-to-text description is performed on the relevant monitoring information by an information processing module;

[0041] (3) A monitoring comparison module performs monitoring comparison based on the analysis results in combination with the relevant monitoring information, generates a monitoring comparison result, and stores it in a monitoring database; and

[0042] (4) Based on the monitoring frequency in the relevant monitoring information, the monitoring comparison results in the monitoring database are brought into a modular template by a report generation module at a specific time interval to generate a monitoring report.

[0043] In this embodiment, step (2) further comprises:

[0044] (21) If the relevant monitoring information input by the user has a trademark name, the name is analyzed by a text parsing unit;

[0045] (22) If the relevant monitoring information input by the user has a trademark graphic, the graphic is analyzed by an image parsing unit and an image-to-text description unit; and

[0046] (23) If the relevant monitoring information input by the user has a text description ofthe graphic, the input text description is semantically understood and analyzed by the text parsing unit.

[0047] In this embodiment, the method of data update and model training by operating a manager electronic device from the manager side, connecting to a server through a first network interface controller of the manager electronic device using a first processor, and executing a first application program, further comprises:

[0048] (5) Extracting trademark public data from multiple trademark data sources by a data extraction module and storing it in an internal database; and

[0049] (6) Training a machine learning model using the large amount of data in the internal database by a model training module and storing the trained data in a model database.

[0050] In this embodiment, step (6) further comprises:

[0051] (61) Text training by the model training module, which is the process of training a machine learning model using trademark name text data;

[0052] (62) Image training by the model training module, which is the process of training a machine learning model using trademark graphic data; and

[0053] (63) Image-to-text description training by the model training module, which is the process of training a machine learning model using images and corresponding text descriptions, enabling the model to understand both the semantics of images and text.

[0054] In this embodiment, graphic data is a data structure used to represent relationships between entities.

[0055] In this embodiment, step (4) further comprises:

[0056] (41) The user adjusts the similarity value settings of a monitoring standard through the input module of the user electronic device based on the content of the monitoring report, thereby increasing or decreasing the number of monitoring comparison results.

[0057] The following provides a detailed explanation using a specific embodiment and diagrams.BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The technical characteristics of this disclosure will become apparent with the detailed description of preferred embodiments accompanied with the illustration of related drawings.

[0059] Figs.1~4 show a schematic diagram of the prior art.

[0060] Figs.5~7 show a schematic diagram of the framework of the system.

[0061] Fig.8 shows a schematic diagram of the example.

[0062] Figs.9~10 show a flow chart of the method.

[0063] Fig.l 1 shows a schematic diagram of an embodiment.

[0064] Fig.12 shows a schematic diagram of another example.

[0065] Fig.13 shows a schematic diagram of another example.DETAILED DESCRIPTION OF THE DISCLOSURE

[0066] One embodiment of the present invention is illustrated in FIGS. 5 to 7, FIGS. 5 to 7 show a basic architecture diagram of the system of the present invention. As shown in FIGS. 5 to 7, the system of the present invention is implemented by providing a user electronic device 200 operated by a user U, wherein the user electronic device 200 may be, but is not limited to: mobile phones, computers, tablets, etc. The user electronic device 200 is connected to a server 400 via a network 300. The user electronic device 200 comprises a second processor 210, a second memory 220, and a second network interface controller 230. The server 400 comprises a second application program 610, so that the second processor 210 is connected to the server 400 via the second network interface controller 230 to execute the second application program 610. At the same time, it is implemented by a manager M operating a manager electronic device 100, wherein the manager electronic device 100 may be, but is not limited to: computers, tablets, etc. The manager electronic device 100 is connected to the server 400 via a network 300. Themanager electronic device 100 comprises a first processor 110, a first memory 120, and a first network interface controller 130. The server 400 comprises a first application program 410, so that the first processor 110 is connected to the server 400 via the first network interface controller 120 to execute the first application program 410.

[0067] Specifically, the first processor 110 is connected to the server 400 via the first network interface controller 120 to execute the first application program 410, thereby configuring and enabling the data extraction module 411, internal database 412, model training module 413, model database 414, and multiple trademark data sources 500 to update data and train models.

[0068] In this embodiment, the inventors used the following server specifications to train and infer the model for this technology: The processor (CPU) is a high-performance multi-core processor, especially for processing large amounts of data and performing complex calculations. At least one CPU with 16 cores or more is used (e.g., AMD Ryzen Threadripper or Intel Xeon series). The memory (RAM) size can be, but is not limited to, 64GB or higher to accommodate the size of the corpus to be processed and the size of the word vector model. The network interface controller is a high-speed and stable network connection hardware, especially for use in cloud computing resources or downloading / uploading large amounts of data. Among them, the most important graphics processing unit (GPU) is a high-performance GPU (such as NVIDIA's RTX 30 series or Tesla series) to reduce the time it takes to train the model. In this case, the training architecture of the server uses a training server with the above high specifications, and when the model training is completed and inference processing is performed, a server host with lower specifications can be used, and it can also be deployed as a cloud or on-premises server host for this system. In this architecture, the use of server host specifications does not affect the technical features emphasized in this case, so any server host specifications should still fall within the technical scope of this case.

[0069] Specifically, manager M operates the first processor 110 of the manager electronic device 100 to execute the data extraction module 411 of the first application program 410. The data extraction module 411 extracts public trademark data from multiple trademark data sources 500 at a regular time interval, such as: trademark names, trademark designs, trademark owners, trademark classes, trademark goods items, countries, etc. The first processor 110 first displays the extracted data on a display screen through the first memory 120 for manager M to confirm, and then stores the extracted data in the internal database 412.

[0070] The first processor 110 further configures and executes the model training module 413. The large amount of data in the internal database 412 is transferred to the model training module 413 for model training. Finally, the trained training data is stored in the model database 414 for subsequent comparison.

[0071] It should be noted that model training and corresponding updates to the model database 414 are required after each data extraction and update.

[0072] Data extraction methods can include, but are not limited to, database APIs (Database APIs), web scraping (Web Scraping), and file-based extraction (File-Based Extraction).

[0073] In detail, database APIs (Database APIs) are interfaces for interaction between applications and databases. They provide a set of standardized functions or methods that allow applications to connect to databases, execute queries, insert and update data, and manage database objects. They can be divided into two types: the first type is Native Database APIs (native database APIs), which are native APIs provided by database vendors and typically provide the most direct and lowest-level access to database functionality. The second type is Third-party Database Connectors (third-party database connectors), which are APIs developed by third-party vendors and typically provide a higher-level abstraction of native APIs and may offer additional features and conveniences.

[0074] The general process can include: establishing a connection, using a database API toestablish a connection to the database, which typically requires specifying connection information for the database, such as hostname, port number, username, and password; executing queries, using the database API to execute SQL queries to extract data, the query can specify the data to be extracted and the extraction conditions; processing results, the database API will return the query results, the application can use these results to update the internal database or perform other processing.

[0075] Specifically, the model training module 413 for the model training of the present invention can be further divided into text training, graphic training, and graphic-to-text description training. The purpose of the model training performed here is to enable the model to be more accurate and refined in subsequent actual trademark monitoring for text, graphics, and graphic-to-text descriptions. For example: it will not recognize the wrong text or mistake a cat graphic for a dog, etc.

[0076] Text-based model training (Text-based Model Training) refers to the process of using text data to train a machine learning model. The training data can be any form of text such as books, articles, news, social media posts, etc. The present invention uses a large amount of Trained on trademark name text, the model uses the training data to learn patterns and rules of the text, which can then be used to perform a variety of tasks. Commonly used steps can include: Data Collection and Preprocessing (Data Collection and Preprocessing), which collects training data and performs preprocessing, such as cleaning data, removing noise, and formatting data; Feature Engineering (Feature Engineering), converting text into a model Understandable numerical features, which usually involves the use of natural language processing (NLP) techniques such as word segmentation, part-of-speech tagging, and word embeddings; Model Selection and Training (Model Selection and Training), choosing an appropriate model architecture and training the model, common models Architecture includes neural networks, such as RNN and LSTM, and Bayesian models; model evaluation (Model Evaluation), which evaluates the performance of the model and fine-tunes it.

[0077] Graph-based Model Training refers to the process of using graph data to train a machine learning model. The present invention uses a large number of trademark graphics for training. The graph data is a data structure used to represent the relationship between entities. . Commonly used steps can include: data collection and preprocessing, which collects graph data and performs preprocessing, such as handling missing values and noise; Graph Representation, which converts graphs into numerical representations that the model can understand, which usually involves using Graph embedding technology; model selection and training, select an appropriate model architecture and train the model. Common model architectures include graph neural network (GNN) and graph convolution network (GCN); model evaluation, evaluate the performance of the model and fine-tune it.

[0078] Model training for converting graphics into text descriptions (Image Captioning Model Training) refers to the process of using images and corresponding text descriptions to train machine learning models. Images can be any form of visual content such as photos, paintings, charts, etc. The text description can be a brief description or a detailed description of the image's content. The model uses image and text description data to learn the relationships between images and text. These relationships can then be used to generate text descriptions of new images. The model can understand the semantics of images and text simultaneously. Commonly used steps can include: data collection and preprocessing, which collects image and text description data and performs preprocessing, such as processing image size, text length, and noise; feature engineering, which converts images and text into numerical features that the model can understand, This usually involves the use of image processing techniques, such as CNNs, and natural language processing techniques, such as word segmentation, part-of-speech tagging, and word embeddings; model selection and training, choosing an appropriate model architecture and training the model. Common model architectures include encoder-decoder models, and attention mechanism models; model evaluation, to evaluate the performance of the model and fine-tune it. Common evaluation indicators include BLEU score, CIDEr score andROUGE score.

[0079] The present invention particularly emphasizes the model training of converting graphics into text descriptions. Through the data acquisition module 411, trademark patterns in multiple trademark data sources 500 from various countries are provided to the model training module 413 for training, so that Each trademark pattern will generate a corresponding text description, and these corresponding text descriptions will be stored in the model database 414 together. Please refer to Figure 8. Figure 8 is an image of a demonstration case. The text description generated based on the image in Figure 8 is: A gray metropolis rises from the ground, and tall buildings cast long shadows under the gloomy sky. , highways wind through concrete jungles, while rivers reflect the sombre tones of the cityscape.

[0080] Specifically, the second processor 210 connects to the server 400 through the second network interface controller 220 and executes the second application 610, thereby configuring and enabling the input module 611, information storage unit 612, information understanding module 613, monitoring comparison module 614, report generation module 615, and monitoring database 700, to achieve trademark monitoring.

[0081] User U operates the second processor 210 of the user electronic device 200 to execute the input module 611 of the second application 610. The input module 611 receives the relevant information that user U wants to perform trademark monitoring and stores it in the information storage unit 612. The input information can include but is not limited to trademark name, trademark image, trademark owner, text description, trademark category, monitoring country, monitoring frequency, monitoring standards, etc.

[0082] The second processor 210 further configures and executes the information understanding module 613. For the input monitoring information, it is analyzed through the text parsing unit 6131, image parsing unit 6132, and image-to-text description unit 6133. And after the above analysis, the results are stored in the information storage unit 612.

[0083] In detail, the text parsing unit 6131 can be regarded as natural language processing, such as LSI (Latent Semantic Indexing). It can include but is not limited to the following methods: Additionally, the comparison method can include, but is not limited to, using string matching algorithms such as Levenshtein Distance and Jaro-Winkler Distance. Levenshtein Distance (also known as Edit Distance) is an algorithm used to measure the similarity between two strings. Its principle is to calculate the minimum number of operations required to transform one string into another, where the operations can be inserting a character, deleting a character, or substituting a character.

[0084] Latent Semantic Indexing (LSI), also known as Latent Semantic Analysis (LSA), is a statistical method that uses singular value decomposition (SVD) to identify patterns between documents and terms. It is a powerful tool for information retrieval and text mining because it can help reveal hidden semantic relationships between words and concepts that may not be discovered by traditional keyword-based methods. It can include but is not limited to the following steps: construct a matrix in which rows represent files, columns represent terms, and each cell in the matrix represents the frequency of a specific term in the corresponding file; apply TF-IDF (term frequency-inverse document frequency), etc. technique to normalize word frequencies across documents, which helps illustrate the importance of terms in individual documents and in the overall corpus; decompose the normalized document-term matrix using SVD, which decomposes the matrix into three matrices: U (left singular matrix), S (diagonal singular value matrix), and VAT (right singular matrix transpose); choose a subset of singular values, usually the top k values, where k is determined by the desired level of dimensionality reduction, which It will reduce the number of potential semantic dimensions while retaining the most important semantic information; the concept space is constructed by multiplying the reduced U matrix by the diagonal matrix of selected singular values. Each row in the resulting concept space matrix represents a file and each column Represent a latent semantic concept; project documents and terms into the concept space using U and VAT matrices respectively,which provide vector representations of documents and terms in a reduced semantic space; calculate cosine similarity between document vectors to measure their semantic similarity.

[0085] The image analysis unit 6132 can be regarded as performing image understanding on the received trademark image, such as labeling the trademark image with feature tags or understanding the basic information of the image, such as shape, size, color, etc., which can include but is not limited to the following methods : Convolutional Neural Network (CNN). CNN is a type of artificial neural network (ANN) that is particularly suitable for image recognition and analysis. They are good at extracting spatial features and patterns from images, making them suitable for tasks such as target detection, image classification and segmentation. Ideal for; deep learning models, such as Recurrent Neural Networks (RNN) and Long Short- Term Memory (LSTM) networks, can be used for tasks involving sequence data, such as image description and visual question answering. These models can capture contextual information and long-term information in images. Distance dependency; transfer learning. Transfer learning involves using a pre-trained model, such as the ImageNet model, to initialize the parameters of a new model. This method can significantly improve the performance of image analysis tasks, especially when dealing with limited training data; attention Mechanisms that allow models to focus on specific regions or features in an image that are most relevant to the task at hand, which can improve the accuracy and efficiency of image analysis, especially in complex scenes or when dealing with multiple objects; feature engineering, which involves extracting data from images Manual extraction and transformation of features, which can be done using techniques such as edge detection, color analysis, and texture analysis. Although less common in deep learning methods, feature engineering can still be used for specific tasks; domain adaptation, when training and testing data come from different When distributing, domain adaptation techniques are used, which may involve techniques such as data augmentation, adversarial learning, and fine-tuning; metric learning, which focuses on learning effective distance measures to compare images or features, which can be used for image retrieval, clustering, andanomaly detection, etc. Task; active learning, which selects the most informative data points to label, thereby reducing the need for large amounts of labeled data, which is particularly beneficial when labeling is expensive or time-consuming; zero-shot learning, which involves classifying images into classes not seen during training , this can be achieved using techniques such as attribute embedding and semantic similarity; explainable artificial intelligence (XAI) technology aims to make image analysis models more interpretable and understandable, which can be achieved using sensitivity analysis, gradient-based method and methods such as Local Interpretable Model Agnostic Explanation (LIME).

[0086] The image-to-text description unit 6133 can be regarded as converting the received trademark image into a text description, that is, texturing the image, and linking the trademark image and the corresponding text description to the model database 414 for comparison. Converting text description is an artificial intelligence technology that aims to convert image content into text description in natural language, or it can be understood as an artificial intelligence technology that allows the system to read or understand the content of the image. Can include but are not limited to the following methods:

[0087] Image feature extraction: First, effective feature information needs to be extracted from the input image. Commonly used image feature extraction methods include: convolutional neural network, which is good at extracting spatial features and patterns from images, and is currently the most commonly used image One of the feature extraction methods; deformer, is a deep learning model based on the attention mechanism that can capture long-distance dependencies in images.

[0088] Language model: The language model is used to generate a text sequence that describes the image content. Commonly used language models include: Recurrent Neural Network (RNN) is a deep learning model that can process sequence data and can adapt to the sequential relationship between words in image descriptions; shapers can also be used aslanguage models, and their parallel processing capabilities allow it to generate textual descriptions more efficiently.

[0089] Alignment of images and text: When converting images to text descriptions, it is necessary to align the extracted image features with the generated text descriptions to ensure that the text descriptions accurately reflect the image content. Commonly used image and text alignment methods include: the attention mechanism allows the model to focus on the image area related to the word when generating each word of the text description; multi-modal fusion, this method can fuse image features and text features together to better learn the alignment relationship between images and text.

[0090] Specifically, the second processor 210 is further configured to execute the monitoring comparison module 614, receives the monitoring information transmitted by the input module 611, and links to the monitoring information combined with the information analyzed by the information understanding module 613. The model database 414 performs monitoring and comparison. For example, the information input by user U is the brand name ABC, the uploaded trademark image, the text description includes an image of the sun and a cat with the cat lying down, and the monitoring country is the United States. , the trademark category is Category 35, the monitoring frequency is every two weeks, the trademark name ABC and text description will be analyzed through the text analysis unit 6131, the trademark image will be analyzed through the image analysis unit 6132 and the image to text description unit 6133, and then The monitoring country, trademark category and monitoring frequency are combined to be linked to the model database 414 for trademark monitoring.

[0091] Since the monitoring country is the United States and the trademark category is Category 35, the monitoring comparison module 614 will first select the data whose application country is the United States and the category is Category 35 in the model database 414, and then perform the comparison. The analysis is performed every two weeks. Finally, after themonitoring comparison, the second processor 210 configures the execution report generation module 615 to generate a monitoring report.

[0092] The monitoring comparison module 614 can be regarded as a comparison of the similarity degree of text, a comparison of the similarity degree of graphics, a comparison of the similarity degree of word vectors, or a comparison of the semantic similarity degree of text descriptions.

[0093] Text similarity comparison aims to quantify the degree of similarity between two or more texts, which involves measuring the lexical, semantic or structural similarity between texts. Can include but are not limited to the following methods:

[0094] Lexical similarity: Lexical similarity focuses on the surface similarity of texts, taking into account the presence or absence of common words or phrases. Common lexical similarity measures include: Jaccard similarity, which calculates the proportion of words shared between two texts; edit distance, which measures the minimum number of edits (insertions, deletions, or substitutions) required to transform one text into another.

[0095] Semantic similarity: Semantic similarity goes beyond lexical similarity and takes into account the meaning and contextual relationship between words. Technologies such as word embeddings are often used: Word2Vec, which generates vector representations of single words based on their distributed semantics; GloVe, a global logarithmic bilinear regression model that learns word vectors. Semantic similarity measures (such as cosine similarity or Euclidean distance) are then applied to compare the vector representations of the text.

[0096] Structural similarity: Structural similarity considers the overall organization and structure of the text, including sentence order, paragraph structure and syntactic relationships. The methods are as follows: Longest Common Subsequence (LCS), which identifies the longest common word sequence between two texts; Edit Distance on Trees (EDT), which measures the minimum number of edits required to convert the syntax trees of two texts.

[0097] Word Embedding is a method of representing words as vectors. These vectors can capture the semantic and syntactic information of words. Word Vector Similarity Comparison refers to the use of word vectors to measure The semantic similarity between two words and the application of word vector approximation in natural language processing, for example:

[0098] Word Sense Disambiguation determines the correct meaning of polysemy words in a sentence. For example, in the sentence "I like to bank money", "bank" can mean "bank" or "river bank", by calculating the word vector The similarity between two words or sentences can be used to determine the meaning of "bank" in the sentence; the semantic similarity calculation (Semantic Similarity Calculation) is used to calculate the semantic similarity between two words or sentences. For example, "king" can be calculated and "queen", or calculate the semantic similarity between "The cat sat on the mat" and "The dog sat on the rug"; information retrieval (Information Retrieval), returns relevant information based on user queries Documents, for example, if the user queries "France", documents related to "France" can be returned, such as "French cuisine", "Paris", and "Eiffel Tower".

[0099] The method of word vector comparison can be but is limited to:

[0100] Cosine Similarity -based Distance Metrics (Cosine Similarity -based Distance Metrics), calculate the cosine similarity between two word vectors, the value is between 0 and 1, 0 means completely dissimilar, 1 means completely Similarity, commonly used distance measures based on cosine similarity include: Cosine Similarity: Cosine Similarity(u, v) = u • v / | |u| | | |v| |; Jaccard Similarity ( Jaccard Similarity): Jaccard Similarity(u, v) = |u v| / |u U v|.

[0101] Euclidean Distance-based Distance Metrics calculates the Euclidean distance between two word vectors. This value represents the distance between the two vectors in the vector space. Commonly used distance metrics based on Euclidean distance include: Euclidean Distance: Euclidean Distance(u, v) = ||u - v||; Manhattan Distance: Manhattan Distance(u , v) = X |ui - vi|.

[0102] Correlation-based Distance Metrics calculate the correlation between two word vectors. This value represents the linear dependence between the two vectors. Commonly used correlation-based distance measures include: Pearson Correlation Coefficient: Pearson Correlation Coefficient^, v) = X (ui - u)(viX (ui - u)2X (vi - v)2.

[0103] Word vector approximation comparison techniques mainly include the following: Database-based Techniques, which store word vectors in the database and use indexes to quickly find similar words; Approximate nearest neighbor-based techniques Nearest Neighbor) technology uses the approximate nearest neighbor algorithm to quickly find the most similar words; Hash Table-based Techniques uses a hash table to quickly find similar words.

[0104] Graph Similarity Comparison is a method used to measure the degree of similarity between two or more graphs. It involves evaluating the structural and semantic similarities between graphs, taking into account node connectivity, Factors such as edge relationships and overall graph structure.

[0105] Structural similarity measures focus on the topological aspects of graphs, emphasizing the arrangement of nodes and edges. Common structural similarity measures include: graph edit distance (GED), which calculates the minimum number of edit operations (insertion, deletion, or edge modification) required to transform one graph into another; subgraph isomorphism, which determines whether a graph contains A subgraph that is isomorphic (structurally identical) to another graph; graph matching, aligning corresponding nodes and edges between two graphs to evaluate their structural similarity.

[0106] Graph edit distance is a measure used to measure the structural similarity between two graphs. It calculates the minimum number of editing operations required to convert one graph into another graph. Editing operations can include inserting nodes and deleting nodes, insert edges, delete edges and replace node or edge attributes. GED is usually calculated using a dynamic programming algorithm, which divides two figures into subgraphs and calculates theminimum number of edit operations required to transform each subgraph into another subgraph. Then, the minimum number of edit operations of these subgraphs is Add the numbers to get the GED of the entire graph.

[0107] Subgraph isomorphism is a method used to determine whether a graph contains a subgraph that is isomorphic (structurally the same) as another graph. If two graphs are isomorphic, they have the same structure, but the nodes and edges may have different labels and properties. Computation of subgraph isomorphisms typically uses graph isomorphism algorithms, which are often based on search techniques such as backtracking or branch-and- bound. Subgraph isomorphism is a powerful graph similarity measure because it can capture an exact match of graph structure.

[0108] Graph matching is a method for aligning corresponding nodes and edges between two graphs. Graph matching can be used to calculate graph similarity, and can also be used for other tasks, such as graph fusion or graph editing. Graph matching is usually calculated using graph matching algorithms, which are often based on techniques such as maximum matching, minimum edit distance, or graph isomorphism.

[0109] Text Description Similarity Comparison refers to a method of measuring the semantic similarity between two or more text descriptions. It is similar to word vector similarity comparison and is mainly used to understand text descriptions. Compared with vectorization, the key technology is to convert the image into the corresponding text description as mentioned above.

[0110] After completing the comparison and analysis based on the monitoring frequency in the input monitoring information, the monitoring comparison module 614 stores the results in the monitoring database 700. The second processor 210 further configures and executes the report generation module 615, and applies the results to the pre-designed template and displays them on the user's electronic device 200 in a display mode such as sorting by similarity andtrademark category. User U will receive the monitoring report according to the monitoring frequency set by himself, and can adjust it at any time.

[0111] Further, the aforementioned monitoring standards can be regarded as adjusting the setting of similarity values, that is, user U can adjust the results that he hopes to monitor to more similar results (less quantity) or more fuzzy results (more quantity) based on the results displayed in the received monitoring report. By adjusting the similarity value parameters, the second processor 210 can configure and execute the monitoring comparison module 614 to compare only the results with higher similarity or the results with lower similarity.

[0112] For another embodiment of the present invention, please refer to and combine Figures 9 to 10. Figures 9 to 10 are the method flowchart of the present invention. As shown in Figures 9 to 10, the steps comprise SI 00 to S400.

[0113] In step S100, administrator M performs data retrieval on multiple national trademark data sources 500 through administrator electronic device 100, network 300, server 400, and data acquisition module 411. For example, trademark name, trademark image, trademark owner, trademark category, trademark goods items, country, etc. data. Further, administrator M can set the time for data acquisition module 411 to automatically perform data acquisition on multiple national trademark data sources 500 within the set time value (time zone) to achieve the effect of automatic update, and the acquired data will be stored in the internal database 412.

[0114] In step S200, the model training module 413 receives the large amount of data stored in the internal database 412 and performs model training, and stores the trained data in the model database 414. It should be noted that steps S100 to S200 can be regarded as a process that is automatically performed regularly within the system.

[0115] In step S300, user U enters the relevant monitoring information for trademark monitoring through the input module 611 of the user electronic device 200. The information entered by the user can include but is not limited to trademark name, trademark image,trademark owner, text description, trademark category, monitoring country, monitoring frequency, monitoring standards, etc. The monitoring comparison module 614 will then perform monitoring comparison based on the input information.

[0116] Specifically, the input module 611 stores the received relevant monitoring information into the information storage unit 612. Then, the information processing module 613 receives the relevant monitoring information from the information storage unit 612 and performs text analysis, image analysis, or image-to-text description. The monitoring comparison module 614 then performs monitoring comparison based on the analysis results and the relevant monitoring information to generate monitoring comparison results.

[0117] In step S400, the report generation module 615 generates monitoring reports by substituting the monitoring comparison results into modular templates according to the monitoring frequency entered by user U at a specific frequency time or time interval, and displays them on the user electronic device 200.

[0118] Further in S400, the user can adjust the setting of the similarity value of the monitoring standard based on the content of the monitoring report. If it is adjusted to a higher standard, the number of monitoring comparison results that may appear in the report may be less. If it is adjusted to a lower standard, the number of monitoring comparison results that may appear in the report may be more. These monitoring reports are stored in the monitoring database 700.

[0119] Further after step S400, if there is at least one similar trademark in the monitoring report, and the user believes that it is sufficient to cause confusion, the user can generate a warning document by selecting the display interface of the user's electronic device 200 to generate a warning document. The report generation module 615 will then substitute the basic information of the similar trademark into a modular template and generate a warning document. The basic information can include but is not limited to the trademark owner, address, agent,trademark category, and goods items.

[0120] Specifically, in step S200, further steps S201 to S203 are included. In step S201, the model training module 413 performs model training for text; in step S202, the model training module 413 performs model training for graphics; and in step S203, the model training module 413 performs model training for image-to-text description.

[0121] Specifically, in step S300, the monitoring comparison module 614 will first determine whether the relevant information for trademark monitoring stored in the information storage unit 612 contains any restriction conditions such as monitoring country, monitoring category, or monitoring type. If so, the monitoring comparison module 614 will select the selected monitoring country, monitoring category, and other restriction conditions from the model database 414 to be monitored, and perform more accurate trademark monitoring. The monitoring type represents whether to monitor word marks, graphical marks, or other types of trademarks.

[0122] Specifically, in step S300, further steps S301 to S303 are included. In step S301, if the information entered by user U has a trademark name, the text parsing unit 6131 will first analyze the input trademark name, and then perform text comparison in the model database 414 through the monitoring comparison module 614. In step S302, if the information entered by user U has a trademark graphic, the image parsing unit 6132 and the image-to-text description unit 6133 will first analyze the input trademark image, and then perform image comparison or text description comparison in the model database 414 through the monitoring comparison module 614. In step S303, if the information entered by user U has a text description of the image, the text parsing unit 6131 will first perform semantic understanding and analysis of the input text description, and then perform text description comparison in the model database 414 through the monitoring comparison module 614. The monitoring comparison results generated in steps S301 to S303 will all be stored in the monitoring database 700. And further in stepS400, the monitoring comparison results will be presented in a visual way by the report generation module 615.

[0123] Specifically, in step S300, if the input is a text description, such as: an image containing an instagram trademark, the text parsing unit 6131 will first perform semantic understanding of this description and understand the meaning of this description, understanding that the instagram trademark is a registered trademark pattern. Then, the monitoring comparison module 614 will receive this message and monitor the image containing the instagram trademark in the model database 414, that is, it will monitor the trademark shown in Figure 13. From another perspective, the technology of the present invention for monitoring and comparing graphics based on text descriptions has the ability to identify registered trademarks through text, and even to identify famous trademarks.

[0124] As can be seen from the above, in step S300, if the input is a text description and the text includes a description of a registered trademark graphic, the monitoring comparison module 614 will first monitor and compare the graphic during the monitoring comparison process, and then continue the monitoring comparison based on other descriptions, such as the spatial relationship or color of the graphic. Also because the present invention can monitor and compare graphics through text descriptions, as can be seen from the above examples, the trademark shown in Figure 13 can be monitored, while in contrast, if only the instagram trademark image is used for identification, the trademark may not be monitored and compared because the instagram trademark occupies a very small proportion.

[0125] Specifically, semantic analysis can perform keyword extraction. Keyword extraction is a natural language processing technique that aims to automatically extract important keywords or phrases from text. Methods can be but are not limited to statistical methods, frequency methods (Frequency-based methods), text statistical methods (Statistical methods), text vectorization methods (Text vectorization methods), or machine learning methods(Machine learning methods).

[0126] Frequency methods judge the importance of a word based on its frequency in the text. Common methods include TF-IDF (Term Frequency-Inverse Document Frequency) and Term Frequency (Term Frequency). TF-IDF considers the frequency of occurrence of a word in the text and its importance in the entire corpus, while Term Frequency only considers the frequency of occurrence of a word in the text.

[0127] Text statistical methods analyze the distribution and correlation of words in text based on statistical models. Common methods include Mutual Information (Mutual Information), Pointwise Mutual Information (Pointwise Mutual Information), and Chi-squared Test (Chi- squared Test). These methods usually need to establish a statistical model between words and text, and calculate the importance of words based on the model.

[0128] Text vectorization methods convert text into vector representations, and then use vector space models (Vector Space Model) to calculate the importance of words. Common methods include Bag-of-Words Model (Bag-of-Words Model), Word Embeddings (Word Embeddings), and text vectorization methods (such as TF-IDF vectorization).

[0129] Machine learning methods use machine learning algorithms to train models to learn the importance of words from text. Common methods include text classification, text clustering, and keyword extraction models. These methods require the use of labeled text data to train the model.

[0130] The above can be implemented by writing a program using a scripting language (such as Python).

[0131] Furthermore, it can include keyword expansion, which is a further function of keyword extraction. It receives the generated keywords and extends them, that is, finds synonyms or related words, and connects to a vocabulary table to compare and generate thesesynonyms.

[0132] For example, WordNet is a synonym generation system based on vocabulary expansion methods, or it uses a large corpus to learn the relationships between words, including co-occurrence and contextual similarity, to generate similarity scores between words, and then generate synonyms or related words. For example, LSI (Latent Semantic Indexing) and LDA (Latent Dirichlet Allocation).

[0133] WordNet is a computerized vocabulary database of English words. It contains a large number of English words and organizes and manages words based on their semantic and syntactic relationships. The purpose of WordNet is to provide a reliable language resource for natural language processing and semantic analysis.

[0134] LSI (Latent Semantic Indexing) is a natural language processing technique used to identify synonyms and related words. Its basic principle is to convert text into a vector space model and then perform singular value decomposition (Singular Value Decomposition, SVD) to identify the semantic relationships between words. Specifically, LSI uses the word frequency statistical method in text to process the words in the text and convert them into a vector space model. Before performing SVD, LSI typically uses TF-IDF (Term Frequency-Inverse Document Frequency) weights to weight the words to eliminate the influence of some common words and better capture the correlation between words. Through such processing, LSI can generate a word-text matrix, where each word corresponds to a vector. By singular value decomposition, LSI can decompose the singular values and singular vectors of this word-text matrix to capture the semantic information contained in the text. Based on this semantic information, LSI can calculate the similarity between two words to find synonyms or related words.

[0135] In addition, LDA (Latent Dirichlet Allocation) is a commonly used topic model that can assign text in a corpus to multiple topics and find the word distribution represented by eachtopic. This word distribution can be used to find synonyms or related words. First, you need to prepare a text collection, which can be any text collection containing the target words. Then, use LDA to assign this text collection to multiple topics and find the word distribution represented by each topic. For a target word, its corresponding word distribution in the LDA model can be used to find synonyms or related words. Specifically, the similarity between the word distribution of the target word and the word distribution of other words can be calculated to find words similar to the target word as synonyms or related words.

[0136] Similarly, a program can be written using a scripting language (such as Python) to execute WordNet, LSI (Latent Semantic Indexing), and LDA (Latent Dirichlet Allocation) to parse text and produce synonyms or related words.

[0137] One application scenario of the present invention is shown in Figure 11, which is described as follows: A user U inputs the trademark information to be monitored through the input module 611 of the user electronic device 200, which is "H2U", the monitoring country is Taiwan, and the monitoring frequency is monthly. This information will be stored in the information storage unit 612. After the information processing module 613 judges that the input information is text, it analyzes "H2U" through the text parsing unit 6131, and then performs monitoring comparison analysis in the model database 414, and stores the results in the monitoring database 700. In the monitoring comparison results presented by the report generation module 615 on the user electronic device 200, the newly applied trademark name "H2O" is listed as highly similar.

[0138] Another application scenario of the present invention is described as follows: A user U inputs the trademark information to be monitored through the input module 611, which is "double x" and the monitoring country is the United States. After the information processing module 613 judges that the input information is a text description, it performs semantic analysis and understanding of "double x" through the text parsing unit 6131, and performs monitoringcomparison analysis on images containing double x and text containing double x in the model database 414, and stores the results in the monitoring database 700. In this application scenario, the presented report will show the trademark of Exxon Mobil, as shown in Figure 12.

[0139] Finally, the technical features of the present invention and the technical effects that can be achieved are summarized as follows:

[0140] First, through a multinational trademark monitoring system and method for semantic understanding of images of the present invention, the problem of traditional manual trademark monitoring is solved, which is time-consuming and labor-intensive, and the results are low.

[0141] Second, through a multinational trademark monitoring system and method for semantic understanding of images of the present invention, the convenience of trademark monitoring for enterprises is improved. It can save the original waiting time and manpower, and users can decide when and what content to receive reports by adjusting the monitoring frequency and monitoring standards.

[0142] Third, through a multinational trademark monitoring system and method for semantic understanding of images of the present invention, through a special model for monitoring graphics with text descriptions, trademarks can be monitored more comprehensively, and the parts that may be missed by traditional manpower or existing technologies can be improved.

[0143] It must be emphasized that the above detailed description is a specific description of an embodiment of the present invention that can be implemented, but this embodiment is not used to limit the scope of the patent of the present invention. Any equivalent implementation or modification that does not deviate from the technical spirit of the present invention should be included in the scope of the patent of the present invention.

Claims

WHAT IS CLAIMED IS:

1. A system of trademark surveillance system with image understanding for multi-country, an user operates a user electronic device, connects to a server through a second network interface controller through a second processor of the user electronic device, and executes a second application program for trademark monitoring, the system comprises at least: an input module for receiving relevant information for trademark monitoring; an information storage unit for storing the relevant information for trademark monitoring received by the input module; an information understanding module for receiving the relevant information for trademark monitoring from the information storage unit and analyzing the information; a monitoring comparison module for performing comparison analysis in a model database based on the relevant information for trademark monitoring in the information storage unit and the text, image, or text description analyzed by the information understanding module to generate a monitoring comparison result; a monitoring database for receiving and storing the monitoring comparison result from the monitoring comparison module; and a report generation module for generating a monitoring report based on the relevant information for trademark monitoring in the information storage unit; wherein, the information understanding module further comprises a text parsing unit, an image parsing unit, and an image-to-text description unit, and performs natural language processing, image understanding, or image-to-text conversion on the relevant information for trademark monitoring.

2. The system according to claim 1, wherein the text parsing unit performs natural language processing on the text in the relevant information for trademark monitoring, vectorizes the text, and stores the analysis results in the information storage unit.

3. The system according to claim 1, wherein the image parsing unit performs image understanding on the images in the relevant information for trademark monitoring, labels the images with feature labels or image basic information, and stores the analysis results in the information storage unit.

4. The system according to claim 1, wherein the image-to-text description unit converts the images in the relevant information for trademark monitoring into descriptive text, texturizes the images, vectorizes the descriptive text, and stores the vectorized descriptive text in the information storage unit.

5. The system according to claim 1, wherein the monitoring comparison module performs text similarity comparison, image similarity comparison, or vector similarity comparison in the model database.

6. The system according to claim 1, wherein the relevant information for trademark monitoring in the information storage unit includes monitoring standards, and the monitoring comparison module compares and analyzes different monitoring comparison results based on the similarity values set by the monitoring standards.

7. The system according to claim 1, wherein an administrator operates an administrator electronic device, connects to the server through a first network interface controller through a first processor of the administrator electronic device, and executes a first application program to update data and train the model, the system comprises at least: a data acquisition module for acquiring trademark public data from multiple trademark data sources; an internal database for storing the trademark public data acquired by the data acquisition module; and a model training module for receiving the trademark public data stored in the internal database and performing model training, and storing the trained training data in the model database;wherein, the data acquisition module acquires data from the trademark data sources based on a frequency setting.

8. The system according to claim 7, wherein the data acquisition module acquires data using database APIs, web scraping, or file-based extraction.

9. The system according to claim 7, wherein the training method of the model training module further comprises text training, image training, and image-to-text description training.

10. The system according to claim 9, the image-to-text description training allows the model to learn the relationship between images and text descriptions using image and text description data, and to use these relationships to generate text descriptions for new images, so that the model can understand both the semantics of images and text.

11. A method of trademark surveillance system with image understanding for multi-country, an user operates a user electronic device to input trademark monitoring related information, connects to a server through a second network interface controller through a second processor of the user electronic device, and executes a second application program, the method of trademark monitoring comprises at least the following steps:(1) a user inputs and stores relevant monitoring information for trademark monitoring using an input module of the user device in an information storage unit;(2) a data processing module performs text analysis, image analysis, or image-to-text description on the relevant monitoring information;(3) a monitoring comparison module performs monitoring comparison based on the analysis results in combination with the relevant monitoring information, generates monitoring comparison results, and stores the monitoring comparison results in a monitoring database;(4) a report generation module generates a monitoring report by bringing the monitoring comparison results from the monitoring database into a modular template using the monitoring frequency in the relevant monitoring information at a specific time interval.

12. The method according to claim 11, wherein the step (2) further comprises:(21) If the relevant monitoring information input by the user has a trademark name, then a text parsing unit analyzes the name;(22) If the relevant monitoring information input by the user has a trademark image, then an image parsing unit and an image-to-text description unit analyze the image; and(23) If the relevant monitoring information input by the user has a text description of the image, then the text parsing unit performs semantic understanding and analysis of the input text description.

13. The method according to claim 11, wherein an administrator electronic device is operated by the administrator, and a first processor of the administrator electronic device is connected to a server through a first network interface controller and executes a first application program to update data and model training further comprises:(5) a data acquisition module acquires trademark public data from multiple trademark data sources and stores it in an internal database;(6) a model training module trains a model using the large amount of data in the internal database and stores the trained data in a model database.

14. The method according to claim 13, wherein the data acquisition module retrieves public data from these trademark data sources through database API, web crawling or file-based extraction.

15. The method according to claim 13, wherein in step (6), it further includes:(61) text training using the model training module is the process of training a machine learning model using trademark name text data;(62) image training using the model training module is the process of training a machine learning model using trademark image data; and(63) image-to-text description training using the model training module is the process oftraining a machine learning model using images and corresponding text descriptions, so that the model can understand both the semantics of images and text; wherein, image data is a type of data structure that represents a digital image.

16. The method according to claim 11, wherein in step (4), it further includes:(41) based on the content of the monitoring report, the user adjusts the approximate value settings of the input module of the user's electronic device to increase or decrease the number of monitoring comparison results.

17. The method according to claim 11, wherein in step (3), it further includes:(311) If the input information includes monitored countries, monitored categories, or monitored types, the monitoring comparison module first performs filtering in the model database.

18. The method according to claim 11, wherein in step (3), it further includes:(321) If the input information contains a trademark name, it is first analyzed by a text parsing unit and then compared to the model database by the monitoring comparison module using text comparison;(322) If the input information contains a trademark image, it is first analyzed by an image parsing unit and then compared to the model database by the monitoring comparison module using image comparison; and(323) If the input information contains a text description of an image, it is first analyzed for semantic meaning by the text parsing unit and then compared to the model database by the monitoring comparison module using text description comparison.

19. The method according to claim 18, wherein in step (3), it further includes:(331) If the input information is a text description of an image, and the content of the description, after being analyzed for semantic meaning by the text parsing unit, describes an image that contains a registered trademark, then the monitoring comparison module first compares the image of the registered trademark in the model database, and then compares thetext of the other descriptive content.

20. The method according to claim 11, wherein after step (4), it further includes:(5) If at least one similar trademark is found in the monitoring report, the user clicks on the generate warning document on the display interface of the user's electronic device, the report generation module then imports the basic information of the similar trademark into the modular template and generates a warning document.

Citation Information

Patent Citations

  • Systems and Methods for Similarity and Context Measures for Trademark and Service Mark Analysis and Repository Searchess

    US20160260033A1

  • Systems and methods for generating text descriptive of digital images

    US20220058340A1