Multi-dimensional vocabulary expansion query system and method for trademark registration risk assessment
By building a multi-dimensional, regionalized trademark registration risk assessment model, using a large language model to generate a set of trademark variants and conduct high-speed search and evaluation, the problems of low efficiency in existing trademark retrieval and inaccurate risk assessment have been solved, and efficient and accurate trademark risk assessment and decision support have been achieved.
Patent Information
- Application Number
- CN202510755411.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-23
AI Technical Summary
Existing trademark search software cannot automatically cover similar forms, sounds, meanings and cross-language variants, resulting in inefficient and error-prone trademark searches, a lack of quantitative risk indicators, high decision-making costs, and the need for multiple rounds of manual searches and comparisons, which takes up a lot of manpower and time.
Build a multi-dimensional, regionalized, quantitative conflict risk assessment and registrability prediction model, use the Large Language Model (LLM) to generate a multi-dimensional text variant set, combine it with the DFA retrieval module for high-speed parallel matching, combine it with regional legal practices to conduct risk assessment, and provide visual reports.
It has achieved improved comprehensiveness of trademark searches and accuracy of risk identification, reduced the risk of missed reports, improved search efficiency, provided quantitative and objective risk assessments, assisted users in making accurate decisions, and improved user experience.
Smart Images

Figure CN120687581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of trademark information retrieval, and in particular to a multidimensional vocabulary expansion query system and method for trademark registration risk assessment. Background Art
[0002] A trademark is the symbol of a company, product, or service. It is integrally linked to the quality of a company's goods, services, and operations, playing a vital role in the industrial and commercial world. It is a key attribute of a company and its products, ensuring uniqueness. To ensure legal protection, a trademark must be officially registered with the Trademark Office.
[0003] With the development of my country's economy and the acceleration of globalization, the number of trademarks has increased year by year. Preventing duplicate or similar registrations is a core issue in trademark management. To protect the legitimate rights and interests of registered trademarks and combat the illegal counterfeiting and misappropriation of registered trademarks, it is necessary to search the pending trademarks and compare them with existing registered trademarks to determine whether they are identical or not similar before they are eligible for registration.
[0004] Due to the continuous growth in the number of trademark registrations and examinations, the trademark database has also grown significantly. Existing trademark search software generally adopts a single-point search mode of "keyword + Nice Classification", which has the following main shortcomings: 1. It can only search the original words entered by the user and cannot automatically cover similar words in form, sound, meaning, and cross-language variants, resulting in omissions.
[0005] 2. To avoid conflicts between similar trademarks, users need to manually conduct multiple rounds of searches and comparisons, which is inefficient and prone to errors.
[0006] 3. Existing software often presents data through paginated result lists, lacks quantitative risk indicators and "registerable" conclusions, and results in high decision-making costs.
[0007] 4. The 45 categories of Nice Classification need to be repeatedly queried and sorted out one by one, which takes up a lot of manpower and time. Summary of the Invention
[0008] In view of this, it is necessary to provide a multidimensional vocabulary expansion query system and method for trademark registration risk assessment that supports regionalized and multidimensional expansion word queries.
[0009] A multi-dimensional vocabulary expansion query system for trademark registration risk assessment, used to build a multi-dimensional, regionalized, quantitative conflict risk assessment and registrability prediction model, including: The user input and parameter configuration module is used to receive the user's query request and configuration parameters, and to verify and encapsulate the input data; the configuration parameters include the name of the trademark to be queried, the category of goods and services to be applied for, and the target market area; The regionalized multi-dimensional vocabulary expansion engine module is used to call the corresponding pre-trained large language model (LLM) based on the target market region of the configuration parameters. In combination with the specific language and cultural characteristics of the target market, it automatically generates a set of text variants of the target vocabulary in multiple dimensions such as form, sound, meaning, colloquialism, and cross-language. DFA search module for high-speed, parallel exact string matching searches of all or a user-specified subset of the Nice Classification in the official trademark databases of the target market area; The regionalized five-dimensional plus penalty risk assessment and registrability prediction module is used to conduct a multi-dimensional quantitative conflict risk assessment on each retrieved prior trademark, combined with the user's proposed target trademark and its extension, based on the legal practices and examination standards of the target market. It then uses the supervised fine-tuned SFT LLM to predict the registrability of the target trademark in the target market region. The result summary and recommendation module is used to generate a structured risk report and list the prior similar trademarks that pose the main risk and detailed conflict information. The risk report includes the target trademark, query parameters, the final registration probability P value, and the inherent success coefficient predicted by LLM; Data visualization and export module, used to display risk assessment reports in a friendly graphical interface and provide report export function; The database and model management module is used to store and manage the regionalized data and models required for system operation, and provide data retrieval and caching services.
[0010] Preferably, the user input and parameter configuration module includes: A user interface interaction unit is used to receive query requests and configuration parameters input by users through a web interface or an API interface; An input verification unit is used to perform basic verification on the length of the entered trademark name, special characters, selected service category, and target market area; Parameter encapsulation unit, used to encapsulate valid user input and configuration parameters into structured data objects; The task distribution unit is used to transfer the encapsulated data object to the regionalized multi-dimensional vocabulary expansion engine module.
[0011] Preferably, the regionalized multidimensional vocabulary expansion engine module includes: An LLM calling and variant generation unit is used to call the selected LLM through the API to obtain a preliminary list of expanded words with regionalized features; The traditional algorithm supplement and regionalized refinement unit is used to generate multi-dimensional vocabulary variants, including form similarity expansion, phonetic similarity expansion, semantic similarity expansion, colloquial expansion, and cross-language expansion, based on the regionalized Prompt algorithm guiding the LLM to conduct multi-dimensional vocabulary expansion around the target trademark; The result deduplication and screening unit is used to screen the variants generated by LLM and supplemented by traditional algorithms based on the initial similarity and relevance with the original word and the language habits of the target market, and to control the scale and quality of the expanded set; The structured output unit is used to construct the final expanded word set into a list of multi-dimensional information including similarity in form, sound, meaning, colloquialisms, and cross-language information, forming an expanded vocabulary set with regional characteristics.
[0012] Preferably, the DFA retrieval module includes: A parallel search execution unit, designed to achieve high-speed retrieval through a deterministic finite automaton (DFA) algorithm, utilizing state machine precompilation and single-scan technology. This unit constructs a finite state machine using a predefined state set Q, a character set Σ, and a state transition function δ. The keyword library is constructed as a Trie tree structure. A single scan can simultaneously detect all predefined keywords, allowing rapid searches for prior trademarks identical to the expanded term within the vast trademark database. The result collection and preliminary filtering unit is used to collect all retrieved prior trademark records and perform preliminary filtering based on application date, trademark status, and target market.
[0013] Preferably, the regionalized five-dimensional + penalty item risk assessment and registrability prediction module includes: An iterative evaluation unit is used to conduct a five-dimensional evaluation of each target trademark or combination of a word extension and a prior trademark based on the target market. The five-dimensional evaluation includes Nice Classification compatibility (C), extension type weight (T), literal / semantic similarity (S), prior trademark strength (St), and the inherent registration success coefficient (LLM_Success_Coeff) predicted by the LLM. The banned word penalty mechanism unit is used to evaluate the target trademark based on the built-in banned word database of the target market area. If the core part of the target trademark hits the banned word list of the target market area, the final registration probability P is forced to be set to an extremely low value; The comprehensive conflict risk coefficient and registrable probability calculation unit is used to calculate the conflict risk coefficient r_single and the final registrable probability P of the target trademark based on the evaluation results obtained by the iterative evaluation unit and the banned word penalty mechanism unit.
[0014] Preferably, the result summary and recommendation module includes: A data aggregation and sorting unit is used to sort prior conflicting trademarks according to the conflict risk coefficient r_single value from high to low, or the final registration probability P value from low to high; The report generation unit is used to fill in the target trademark information, query parameters, P value, LLM_Success_Coeff, and high-risk trademark list according to the preset template. The high-risk trademark list includes the name, application number, category, Nice Classification compatibility C, extended type weight T, literal / semantic similarity S, prior trademark strength St, conflict risk coefficient r_single, and registration country / region.
[0015] Preferably, the data visualization and export module includes: A visualization unit, which uses a front-end chart library to visually display the report data generated by the result summary and recommendation module in the form of a dashboard, risk list, or regional risk heat map, and indicates the target market area. If multi-market evaluation is supported, it provides switching or parallel display of evaluation results; The report export unit is used to export the complete risk assessment report including all charts and detailed data into the standard PDF format with one click.
[0016] Preferably, the database and model management module includes: Regionalized trademark databases, used to store and regularly update official trademark data for target market regions, including trademark names, images, applications / registrations, applicants, designated goods / services classifications, legal status, application / registration dates, and trademark strength information; Regionalized model library, used to store pre-trained LLM models and regionalized LLM models after SFT; User database, used to store user information, query history, and configuration preferences; A regionalized banned word database, which stores banned word lists, legal basis, and severity levels for each target market region; Regionalized rules and parameter library, used to store the Nice Classification compatibility rules, extended type weights, and trademark strength calculation rules for each target market region.
[0017] A multi-dimensional vocabulary expansion query method for trademark registration risk assessment is provided. The multi-dimensional vocabulary expansion query system for trademark registration risk assessment provides users with a trademark registration risk assessment service for a specific market. The method includes the following steps: Step 1: The user initiates a query request and enters configuration parameters; Step 2: Based on the target market region selected by the user, the corresponding pre-trained LLM is called. Combining the regional language and cultural characteristics, the target trademark is expanded in multiple dimensions, including form, sound, meaning, and across languages, to generate an expanded vocabulary set containing multiple potential close variants. Step 3: Aggregate the target vocabulary input by the user and the extended vocabulary set to form a target extended vocabulary set; Step 4: Perform a high-speed parallel search of the official trademark database corresponding to the target market area for the user-specified category and related similar categories; Step 5: Determine whether a prior similar trademark is found. If no prior similar trademark is found, proceed to Step 8. If a prior similar trademark is found, proceed to Step 6. Step 6: For each similar trademark retrieved, conduct a detailed risk assessment in combination with the target trademark or the expanded vocabulary set, and calculate the individual risk coefficient of each potential conflict. , and finally summarize the maximum conflict risk and the comprehensive registration probability P; Step 7: After all prior trademarks have been evaluated, the analysis results are summarized and a structured risk report is generated. Based on the evaluation results, a risk assessment is conducted to determine whether there is a risk. If there is a risk, the high-risk prior trademarks and their conflict details, as well as the inherent registrability risk, are listed. If there is no significant conflict, only the inherent registrability risk is displayed. Step 8: Generate a structured risk report; Step 9: Data visualization and export; the generated report content is displayed on the user interface in visual forms such as charts, and the function of exporting the complete report to PDF format is provided.
[0018] Preferably, in step 6, for each retrieved prior similar trademark, a detailed risk assessment is performed in combination with the target trademark or the expanded vocabulary set, and the individual risk coefficient of each potential conflict is calculated. , and finally summarize the maximum conflict risk The steps of calculating the comprehensive registration probability P specifically include: Step 6.1, receiving input parameter values, including target word, previous word, target market area, product / service category, and target word expansion type; Step 6.2: Calculate the Nice Classification compatibility, C. Based on the "Classification of Similar Goods and Services" for the target market area or case law, determine whether the target goods and the goods with the prior trademark belong to the same or similar categories. A comprehensive assessment of similarity is made based on the nature, use, and distribution of the goods / services, and the score is mapped to a similarity score. The value of C ranges from 0 to 1. Step 6.3: Determine the weight T of the expansion type. When determining the impact of each dimension of lexical expansion on the likelihood of confusion, T ranges from 0 to 1. The weight T can be fine-tuned based on regional legal practice, giving different weights to similar-sounding expansions, similar-sounding expansions, regionalized transcriptions, and similar-sense expansions. Step 6.4, calculate the literal / semantic similarity S; calculate the similarity between the target trademark or its extension and the prior trademark name, including: Similar extension: The normalized Levenshtein distance is shown in formula (1): Sortho=1−(LevenshteinDistance / max(len(str1),len(str2)))(1); Among them, Sortho represents the normalized literal similarity, and the value range of Sortho is 0~1; LevenshteinDistance represents the edit distance between two strings, which is the minimum number of single-character edits required to convert one string to another; len(str1) / len(str2) represents the length of the first string and the second string respectively; Phonetic extension: Determines the degree of pronunciation similarity between the target trademark or extension and the prior trademark name. The pronunciation similarity Sphonetic value ranges from 0 to 1. When the regionalized phonetic code is a perfect match, Sphonetic = 1.0; Similarity expansion, which is the semantic similarity between the target word and the original word given when LLM is used for vocabulary expansion, or the semantic similarity between the target word and the prior trademark calculated through regionalized LLM; Step 6.5: Evaluate the strength of the prior trademark (St); determine the strength of the prior trademark's distinctiveness, reputation, and legal protection in the country of its registration; Step 6.6: Obtain the inherent registration success coefficient LLM_Success_Coeff; the large language model (SFT) fine-tuned by the official trademark review opinions in the target market region is used to directly predict the inherent registrability and initial success probability of the target trademark in the corresponding regional market; Step 6.7, check for banned words and build a banned word list for the target market area; Step 6.8, determine whether the target word hits the banned word; if the target trademark hits the banned word in the target market area, then, =0.01, or a more detailed penalty value determined by the severity of the banned words, otherwise, = ; Step 6.9, calculate the single conflict risk coefficient The single conflict risk coefficient is the degree of conflict between the target trademark or expansion word and each prior trademark. The calculation formula is shown in formula (2): (2); in, 、 、 are the Nice Classification compatibility C, the extension type weight T, and the literal / semantic similarity S weights, respectively. 、 、 The sum of the three is 1; is the strength of the prior trademark, The value range is 0~1; Step 6.10, output intermediate results; output the current combination of Nice Classification compatibility C, extended type weight T, literal / semantic similarity S, prior trademark strength St, and single conflict risk coefficient Value, as well as the inherent registration success coefficient LLM_Success_Coeff, and Bad Words hit situation. Among them, the inherent registration success coefficient LLM_Success_Coeff is fixed for the same target trademark in the same target market area query; Step 6.11: End the single risk assessment. Repeat steps 6.2 to 6.10 for each combination of the target trademark or word extension and each prior trademark until all retrieved prior trademarks have been assessed. Step 6.12, Determine the Maximum Conflict Risk and the comprehensive registration probability P; 1) Maximum risk of conflict : is the conflict coefficient of all single lines The maximum value among the above is obtained by comparing the target trademark with all prior trademarks. The maximum value of is calculated as shown in formula (3): = max( , ,..., )(3); If there is no conflict, then =0; 2) The final registration probability P includes: i) Basic probability , the calculation formula is shown in formula (4): =LLM_Success_Coeff(4); Among them, LLM_Success_Coeff is the inherent registration success coefficient predicted by LLM; ii) Conflict-adjusted probability , the calculation formula is shown in formula (5): = ×(1− )(5); Among them, 1− ≥0; if ≥1 then 1− Considered as 0; iii) Banned word penalty: If the target trademark hits the banned words in the target market area, then, =0.01, or a more detailed penalty value determined by the severity of the banned words, otherwise, = ; Step 6.13: Calculate the conflict risk coefficient r_single and the final registration probability P of the target trademark and transmit them to the result summary and recommendation module: Among them, the value of the conflict risk coefficient r_single is the maximum conflict risk ; The final registration probability P is value.
[0019] The aforementioned multidimensional vocabulary expansion query system and method for trademark registration risk assessment includes a user input and parameter configuration module, a regionalized multidimensional vocabulary expansion engine module, a DFA search module, a regionalized five-dimensional plus penalty risk assessment and registrability prediction module, a result aggregation and recommendation module, a data visualization and export module, and a database and model management module. By constructing a multidimensional, regionalized, quantitative conflict risk assessment and registrability prediction model, leveraging advanced large language models and incorporating regional linguistic and cultural characteristics, it intelligently generates a comprehensive set of variants of the target trademark in multiple dimensions, including form, pronunciation, meaning, colloquialisms, internet slang, and cross-language transliterations. The system automatically completes the entire process of vocabulary expansion, full-category parallel search, and risk calculation and prediction. The system boasts excellent scalability and maintainability, significantly improving the comprehensiveness of trademark searches and the accuracy of risk identification, effectively reducing the risk of underreporting. It also enhances the efficiency of trademark searches and risk assessments, enabling "one-click" rapid response. It provides quantitative, objective, and regionally adapted risk assessments and registrability predictions to assist users in making accurate decisions. Furthermore, it improves the user experience with intuitive visual reports and convenient export capabilities. The method of the present invention is simple, easy to implement, low in cost and easy to promote. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a structural diagram of a multi-dimensional vocabulary expansion query system for trademark registration risk assessment according to an embodiment of the present invention.
[0021] Figure 2 It is a flowchart of a multi-dimensional vocabulary expansion query method for trademark registration risk assessment according to an embodiment of the present invention.
[0022] Figure 3 This is a flowchart of the conflict risk scoring of the multidimensional vocabulary expansion query method for trademark registration risk assessment according to an embodiment of the present invention. DETAILED DESCRIPTION This embodiment takes a multi-dimensional vocabulary expansion query method, system, and computer-readable storage medium for trademark registration risk assessment as an example, and the present invention will be described in detail below in conjunction with specific embodiments and drawings.
[0023] Example 1 See also Figure 1 This article illustrates a multi-dimensional vocabulary expansion query system for trademark registration risk assessment, provided by an embodiment of the present invention. This system is used to construct a multi-dimensional, regionalized, quantitative conflict risk assessment and registrability prediction model. This system is designed to be deployed on standard server hardware or mainstream cloud computing platforms. Recommended hardware configurations include: CPU: Multi-core processor (e.g., 8 cores or more, and GPU can be considered for LLM reasoning); Memory: 32GB RAM or above (adjusted based on the amount of data processed, LLM model size, and number of concurrent users); Storage: High-speed SSD storage, the capacity depends on the size of the trademark database, model files and log requirements; Network: High-speed network connection for API calls, database access, and user interaction.
[0024] The multi-dimensional vocabulary expansion query system for trademark registration risk assessment includes: 1. User input and parameter configuration module 10, used to implement the following functions: 1. Receive the trademark name (text) to be queried entered by the user.
[0025] 2. Allows users to specify the categories of goods / services for which the trademark is to be applied for (for example, by selecting the Nice Classification symbol).
[0026] 3. Allow users to select the target market / country (currently focusing on China and the United States), and the system will call the corresponding regionalized models, rules and databases accordingly.
[0027] 4. Allow advanced users to adjust the sensitivity or weight of some risk assessment parameters, provide default configurations in the user interface, and explicitly indicate the risk of adjustment.
[0028] The processing steps of the user input and parameter configuration module 10 are as follows: 1. User interface interaction: Receive user input through the web interface or API interface.
[0029] 2. Input verification: Perform basic verification on the length of the entered trademark name, special characters, selected category and market.
[0030] 3. Parameter encapsulation: Encapsulate valid user input and configuration parameters (especially target market) into structured data objects.
[0031] 4. Task distribution: The encapsulated data object is passed to the regionalized multidimensional vocabulary expansion engine module to start the processing flow.
[0032] 2. Regionalized multi-dimensional vocabulary expansion engine module 20 is used to implement the following functions: 1. Based on the target market (China / US) selected by the user in the User Input and Parameter Configuration module, the corresponding pre-trained Large Language Model (LLM) is called. For example, if the target market is "China", the Qwen3 series model is used; if the target market is "US", the Gemma3 series model is used.
[0033] 2. Utilizing the selected LLM and combining it with the specific linguistic and cultural characteristics of the target market (e.g., Chinese idioms, internet slang, homophonic conventions; American slang, common spelling variations, etc.), the system automatically and intelligently generates a collection of text variants of the target vocabulary in multiple dimensions, including similarity in form, sound (including possible variants influenced by dialects), meaning (including metaphors and extended meanings), colloquialisms, and cross-language (e.g., Chinese-English translation, and common transliteration conventions for the target market).
[0034] The processing steps of the regionalized multidimensional vocabulary expansion engine module 20 are as follows: 1. Receive input: Receive parameter objects containing information such as target trademark, target market area, and service category from the user input and parameter configuration module.
[0035] 2. Regionalized LLM selection: Dynamically load or call the corresponding LLM based on the "target market" field in the parameter object.
[0036] 3. Regionalized Prompt Engineering: Build prompts optimized for target markets and language characteristics. Prompts will guide LLM in developing multi-dimensional, high-quality vocabulary expansions around the target trademark. These may include specific requirements for the type of expansion, expected number, relevance, and culturally specific variations.
[0037] 4. LLM call and variant generation: Call the selected LLM through the API to obtain a preliminary list of expanded words with regionalized features.
[0038] 5. Traditional algorithm supplementation and regionalization refinement (can be used in parallel with LLM or as post-processing, and guided by regionalization rules): 1) Similar to extension: Algorithm: Levenshtein distance.
[0039] Formula meaning: Calculates the minimum number of single-character edits (insertion, deletion, or replacement) required to convert two strings from one to another.
[0040] Parameter: The edit distance threshold d ≤ 1 (which can be fine-tuned according to the language characteristics of the target market, such as the complexity of Chinese characters).
[0041] Application scenario: To capture minor spelling mistakes or single-character variants.
[0042] Example: For the target word "APPLE" (in the US market), with d ≤ 1, it can be extended to "APPLES", "APLE", "XPPLE", etc. For the target word "华为" (in the Chinese market), minor variants with similar strokes may be considered.
[0043] 2) Phonetic similarity extension (regionalized processing): Algorithm: For English, the Double-Metaphone algorithm is used; for Chinese, it is first converted to pinyin (such as using the Pypinyin library, considering the handling of polyphonic characters), and then phonetically similar matching can be performed based on pinyin or phonetic similarity rules for specific regions (such as the influence of dialects).
[0044] Meaning of the formula: Double-Metaphone converts words into phonetic codes, and words with similar pronunciations tend to generate the same or similar phonetic codes.
[0045] Application scenario: To identify words with the same or similar pronunciations, taking into account the confusion that may be caused by regional accents.
[0046] Example: The Double-Metaphone codes of the English words "Smith" and "Smyth" are the same. The pinyin of the Chinese words "快客" (kuài kè) and "哙客" (kuài kè) is the same.
[0047] 3) Semantic similarity extension (mainly covered by LLM, can be used as verification or supplement for specific scenarios): Algorithm: Replacement based on synonyms, near synonyms, and antonyms (in some cases, antonyms may also cause associations) in a specific domain dictionary of the target market.
[0048] 4) Cross-lingual / transliteration extension (regionalized processing): Algorithm: For common translations (official or folk) and transliteration habits in the target market. For example, common English translations or standard pinyin representations of Chinese trademarks; common Chinese transliterations of English trademarks (considering differences in translation methods in different historical periods or regions).
[0049] Example: The pinyin of "华为" is "HUAWEI"; the common Chinese transliteration of "Google" in the Chinese market is "谷歌".
[0050] 6. Result Deduplication and Screening: Deduplication is performed on variants generated by LLM and supplemented by traditional algorithms. This can be done based on initial similarity to the original word (determined by the LLM or algorithm), relevance, and whether it conforms to the target market's language habits, thereby controlling the size and quality of the expanded set.
[0051] 7. Structured output: The final expanded word set is constructed into a list containing information such as {expanded_token, original_token, expansion_type (e.g., LLM_meaning_CN, Algo_shape_US), source_engine (LLM / Algo), market, preliminary_score (e.g., LLM confidence, Levenshtein scorenormalized)}.
[0052] 8. Data transfer: Pass the structured expanded word list to the DFA retrieval module.
[0053] 3. DFA retrieval module 30 is used to implement the following functions: 1. Receive the user input and the extended vocabulary set with regional characteristics generated by the parameter configuration module.
[0054] 2. High-speed, parallel, exact string matching searches of all 45 Nice Classifications (or user-specified subsets) in the official trademark databases (or their mirrors) corresponding to the target markets.
[0055] The processing steps of the DFA retrieval module 30 are as follows: 1. Receive an expanded word set: obtain an expanded word list from the user input and parameter configuration module.
[0056] 2. Construct the search query: Use semicolons to connect each expanded word to form a whole paragraph of text.
[0057] 3. Parallel search execution: Indexing strategy: Maintain independent trademark database indexes optimized by Nice Classification or major class combinations for each target market (China / US).
[0058] 1) Algorithm Application: The search engine utilizes the principles of Deterministic Finite Automata (DFA) for efficient and precise string matching. Using pre-built state transition graphs, DFA can determine whether one or more pattern strings exist in the input text in linear time.
[0059] 2) Formula: A DFA recognizes patterns through a series of states and state transitions based on the input characters. Given a set of patterns (in this case, the names of prior trademarks), a DFA can be constructed that recognizes any of these patterns.
[0060] 3) Application Scenario: Quickly search for prior trademarks identical to the expanded term within a massive trademark database. Searches requiring prefixes, suffixes, or inclusion relationships can also be implemented through appropriate DFA construction or integration with other search engine capabilities.
[0061] 4) Example: Build a DFA (or equivalent index structure) for all prior trademark names, then run each expansion word as input on the DFA to determine if there is an exact match.
[0062] 4. Result collection and preliminary filtering: Collect all retrieved prior trademark records and perform preliminary filtering based on application date, trademark status (such as valid, invalid, pending), target market, etc.
[0063] 5. Data transfer: The retrieved relevant prior trademark dataset (including information such as trademark name, application number, registration number, applicant, designated goods / services, current status, country / region of application, etc.) is transferred to the regionalized five-dimensional + penalty item risk assessment and registrability prediction module.
[0064] 4. Regionalized five-dimensional + penalty item risk assessment and registrability prediction module 40 is used to achieve the following functions: This module is the intelligent core of the system. It combines each prior trademark retrieved by the DFA Quick Search Module with the user's proposed target trademark (and its related extensions) to conduct a multi-dimensional quantitative conflict risk assessment based on the legal practices and examination standards of the target market. It also uses the LLM verified by SFT to predict the registrability of the target trademark in that market.
[0065] The processing steps of the regionalized five-dimensional + penalty item risk assessment and registrability prediction module 40 are as follows: 1. Receiving data: receiving prior trademark data from the DFA retrieval module, and obtaining target trademark information (including its original form and selected target market) from the user input and parameter configuration module and the regionalized multi-dimensional vocabulary expansion engine module.
[0066] 2. Iterative Evaluation: For each combination of "target trademark (or one of its extensions) vs. prior trademark," the following five-dimensional evaluation is conducted based on the target market: a. Nice Classification compatibility (C) (regionalization): Rules: Based on the target market's "Classification of Similar Goods and Services" (such as China's official version) or case law (such as the USPTO's TMEP guidance and precedents).
[0067] Example values (adjusted for market): The target product and the prior trademark product belong to the same official similarity group (e.g., China): C = 1.0.
[0068] Belonging to different similar groups but constituting similar goods in the legal sense (such as Chinese case law): C = 0.7.
[0069] In the US market, similarity is determined based on a comprehensive assessment of the nature, purpose, and channel of goods / services, and can be mapped to a similarity score.
[0070] For distant or dissimilar goods / services: C = 0.3 (or lower, or even 0).
[0071] For example: If the target trademark is used for “clothing” (Class 25, Chinese market) and the prior trademark is used for “ties” (Class 25, same similar group), then C=1.0.
[0072] b. Extended type weight (T): Rules: Reflect the extent to which different types of similarity (determined by the expansion_type of the regionalized multidimensional vocabulary expansion engine module) contribute to the likelihood of confusion, with weights that can be fine-tuned based on regionalized legal practices.
[0073] Numerical examples (basic values, can be adjusted regionally): Phonetic: T = 1.0 Orthographic: T = 0.9 Regionalized Transliteration: T = 0.8 Semantic (via LLM): T = 0.6 For example, if the target trademark’s word extension “APPY” is derived from “APPLE” through similarity in form and is compared with the prior trademark “APPYMORE” (in the US market), then T=0.9.
[0074] c. Literal / Semantic Similarity (S) (0-1 range): Rule: Comprehensively calculate the similarity between the target trademark (or its extension) and the prior trademark name.
[0075] Calculation method (depending on the extension type and market selected): For shape similarity: normalized Levenshtein distance. Formula: Sortho = 1 − (LevenshteinDistance / max(len(str1), len(str2))).
[0076] For phonetic similarity: If the regionalized phonetic code (such as English Double-Metaphone code, Chinese Pinyin sequence) matches completely, then Sphonetic = 1.0; if it matches partially or based on phonetic similarity rules, different values can be set.
[0077] For semantic similarity: the semantic similarity with the original word given by LLM in the expansion stage can be directly used, or the target word and the prior trademark can be again calculated with cosine similarity through a regionalized LLM (such as the regionalized version of Sentence-BERT).
[0078] Numerical example: The target word "SOLUTION" is compared with the prior trademark "SOLUSHUN" (US market). The Levenshtein distance is 1 and max_len is 8. Then Sortho = 1 − (1 / 8) = 0.875.
[0079] d. Strength of prior trademark (St) (regionalized, 0-1 range): Rule: Assess the distinctiveness, fame and strength of legal protection of the prior trademark in the country where it is registered.
[0080] i) Considerations and numerical examples for the US market (St_US): Obtaining §15 Affidavit filed: +0.3 Multiple successful rights protection records (external data support required): +0.2 Made-up or arbitrary word (strong inherent salience): +0.2 High market visibility (based on evidence such as sales, advertising investment, etc., requires external data or user input): +0.1 to +0.3 The basic score is 0.1, and the sum of all the points does not exceed 1.0.
[0081] For example, if a US trademark is a made-up word and is protected under Article 15, St_US = 0.1 (base) + 0.3 (Article 15) + 0.2 (made-up) = 0.6.
[0082] ii) Chinese Market (St_CN) Considerations and Numerical Examples: Recognized as a well-known trademark (official or judicial recognition): St_CN = 0.8 to 1.0 High originality (such as made-up words): +0.2 Long market usage time, wide range, and high reputation (based on usage evidence): +0.1 to +0.3 Having a record of winning in prior right conflicts (such as successful opposition or invalidation): +0.1 The base score is 0.1, and each item is added up, with the maximum not exceeding 1.0.
[0083] Example: A certain Chinese trademark is a coined word with a relatively high market reputation. St_CN = 0.1 (base) + 0.2 (coined) + 0.2 (reputation) = 0.5.
[0084] e. Registered success coefficient predicted by LLM (LLM_Success_Coeff) (regionalized, in the range of 0 - 1): Rule: The large language model (such as Gemma3 - SFT - US / Qwen3 - SFT - CN) supervised fine - tuned (SFT) by the official trademark review opinions in a specific region (such as Office Actions of the US USPTO, examination decisions / notice of rejection of the CNIPA in China) directly predicts the inherent registrability and initial success probability of the target trademark in this market. This coefficient mainly reflects the distinctiveness of the trademark itself, whether it is descriptive, whether it conforms to public order and good customs, etc., and does not directly consider the conflict with a specific prior trademark (the conflict is handled by the other four dimensions).
[0085] Features input to the SFT LLM: Target trademark name, designated goods / service categories, target market.
[0086] Output: A probability value from 0 to 1, representing the "registrability potential" of the trademark itself.
[0087] Numerical example: The Qwen3 - SFT - CN model predicts that the inherent registered success coefficient of a trademark "Zhixingzhe" in the category of robots designated in China is 0.65.
[0088] 3. "Bad Words" penalty mechanism (regionalized):[[ID=2
[0090] For example, if the target trademark is "Great Hall of the People" and is searched in the Chinese market, and it is a proprietary name and contains banned words, then P is forced to be 0.01.
[0091] 4. Calculation of the comprehensive conflict risk coefficient r_single and the registration probability P: i) Single conflict coefficient : The degree of conflict between the target trademark (or one of its extensions) and a prior trademark.
[0092] Formula example: ,in , , is the weight of each factor, for example, they can all be set to 1 / 3, or learned based on a large amount of case data, and the total is 1. As a multiplying factor, the greater risk of prior trademarks is emphasized.
[0093] Numerical example: _US=0.6, C=1.0 (similar), T=0.9 (similar in form), S=0.875 (similar in wording). Assume that the weights are all 1 / 3. =0.6×((1 / 3)×1.0+(1 / 3)×0.9+(1 / 3)×0.875)=0.6×(0.333+0.3+0.2917)≈0.6×0.9247≈0.555.
[0094] ii) Maximum risk of conflict : Among all the prior trademarks compared with the target trademark, the largest value. =max( , ,..., ), if there is no conflict, then =0.
[0095] iii) Final registration probability P: Base probability (inherent registrability): =LLM_Success_Coeff .
[0096] Conflict-Adjusted Probability (Relative Registrability): = ×(1− ), where 1− Not negative, if ≥1 then 1− Treated as 0.
[0097] Applying a banned word penalty: If (the target trademark hits the banned words in this area) Then =0.01 (or a more detailed penalty value determined by the severity of the banned words) Else = .
[0098] Numerical example: If LLM_Success_Coeff = 0.65, = 0.555, and no banned words are hit, we can get =0.65×(1−0.555)=0.65×0.445≈0.289, that is =0.289 indicates that the risk of successful registration due to conflict with prior trademarks is high.
[0099] Specifically, The higher the value of , the higher the probability of registration and the lower the risk; conversely, the lower the Pfinal value, the lower the probability of registration and the higher the risk.
[0100] 5. Data transfer: The conflict analysis details (C, T, S, St, r_single) of each prior trademark, as well as the final P value and LLM_Success_Coeff and other information, together with the target market identifier, are transferred to the result summary and recommendation module.
[0101] 5. Result summary and recommendation module 50 is used to implement the following functions: 1. Summarize all regionalized risk assessment results output by the regionalized five-dimensional + penalty item risk assessment and registrability prediction module.
[0102] 2. Generate a structured risk report that clearly lists the target trademark, query parameters (including target market), the final probability of registration P value, and the inherent success coefficient of LLM prediction.
[0103] 3. List the prior similar trademarks that pose the main risks and their detailed conflict information (differentiate the market).
[0104] 4. (Optional) Based on risk analysis, provide preliminary trademark modification suggestions or alternative word search prompts, and the suggestions may take into account regional culture and word preferences.
[0105] The processing steps of the result aggregation and recommendation module 50 are as follows: 1. Receive assessment data: Obtain detailed risk assessment results with regional identification from the regionalized five-dimensional + penalty item risk assessment and registrability prediction module.
[0106] 2. Data aggregation and sorting: Sort prior conflicting trademarks by risk level (e.g., conflict risk coefficient r_single value from high to low, or final registrable probability P value from low to high).
[0107] 3. Report generation: Based on the preset template, fill in the target trademark information, query parameters, final registration probability P value, inherent registration success coefficient LLM_Success_Coeff, and high-risk trademark list (including their name, application number, category, Nice Classification compatibility C, extension type weight T, literal / semantic similarity S, prior trademark strength St, conflict risk coefficient r_single, registration country / region, etc.).
[0108] 4. Data transfer: pass the generated report data and summary information to the data visualization and export module.
[0109] 6. Data visualization and export module 60 is used to implement the following functions: The risk assessment report is displayed in a user-friendly graphical interface and the report export function is provided.
[0110] The processing steps of the result aggregation and recommendation module 60 are as follows: 1. Visual presentation: Utilize front-end chart libraries (such as ECharts, D3.js) to visually display the summary of the results and the report data of the recommendation module in the form of a dashboard (displaying P-value and LLM success coefficient), a risk list (details of high-risk trademarks), and (optional) a regionalized risk heat map.
[0111] 2. Regionalized Results: Clearly indicate on the interface which target market the currently displayed evaluation results are for. If multiple markets can be evaluated simultaneously, a switching or parallel display function should be provided.
[0112] 3. Report export: Supports one-click export of complete risk assessment reports including all charts and detailed data into standard formats such as PDF. The target market of the assessment should be clearly marked in the report.
[0113] 4. Mutual Relationship: The data visualization and export module receives report data from the result summary and recommendation module, and presents the final results to the user or provides the user with downloading.
[0114] 7. Database and model management module 70, used to implement the following functions: 1. Store and manage various regionalized data and models required for system operation.
[0115] 2. Provide data retrieval and caching services to improve system performance and regional response capabilities.
[0116] The database and model management module 70 includes: 1. Regionalized trademark database: This database stores official trademark data for China and the United States (and potentially other countries / regions in the future). This data includes trademark name, image (if available), application / registration number, applicant, designated goods / services (Nice Classification), legal status, application / registration date, and trademark strength information (such as US Section 15 status and China Well-Known Trademark recognition records).
[0117] Data sources: official public data from various countries and authorized commercial data providers.
[0118] Update mechanism: Regularly (weekly) perform incremental or full updates on each regional database.
[0119] Storage technology: PostgreSQL database (with independent indexing strategies configurable for different regions), supporting efficient full-text search and complex queries.
[0120] 2. Regionalized model library: Stores pre-trained LLM models (Qwen3 series, Gemma3 series).
[0121] Stores the SFT regionalized LLM models (Qwen3-SFT-CN, Gemma3-SFT-US).
[0122] Storage method: Model files are stored in the file system or object storage, and the database records model metadata, version, applicable area, and path.
[0123] 3. User database: stores user information, query history (including target market), configuration preferences, etc.
[0124] 4. Regionalized "banned word" database: stores banned word lists, their legal basis and severity levels by region (China, the United States, etc.).
[0125] 5. Regionalized rules and parameter library: stores the Nice Classification compatibility rules, extended type weights, trademark strength calculation rules, etc. for different markets.
[0126] 6. Cache system: Technology: Redis Cached content: Regionalized LLM expansion results for frequently searched queries, frequently searched prior trademark data, and calculated regionalized risk assessment results (for identical query inputs and markets) to reduce repeated calculations and database accesses and improve response speed.
[0127] Application examples: Assume that the user input parameters are: Target trademark: "Kezhilian"; Designated Class: Class 9 (Scientific Instruments, Computer Software, etc.); Target Market: China.
[0128] 1. The user input and parameter configuration module receives and validates the input, and transmits the target market "China" and other information to the regionalized multi-dimensional vocabulary expansion engine module.
[0129] 2. The regionalized multi-dimensional vocabulary expansion engine module selects the Qwen3 model, expands the vocabulary for "Kezhilian" in combination with Chinese language and culture, and may generate: Shape similarity (considering Chinese character structure): "Kezhilian", "Kezhilian" Phonetic similarity (pinyin kē zhì lián): "Kezhilian", "Kezhilian" Semantic similarity (combining the meanings of technology, intelligence, and connection): "Zhilian Technology", "Intelligent Interconnection" Output a structured list, such as [{token: "Kezhilian", type: "LLM_Phonetic similarity_CN",...}, {token: "Intelligent Interconnection", type: "LLM_Semantic similarity_CN",...}].
[0130] 3. The DFA retrieval module retrieves these expanded words in parallel in the trademark databases of Class 9 and similar classes in China. Assume that the prior trademark "Kezhilian" (Class 9, registered) is retrieved.
[0131] 4. The regionalized five-dimensional + penalty term risk assessment and registrability prediction module assesses "Kezhilian" (or its expanded word "Kezhilian") vs "Kezhilian" (China market): C (Category Compatibility): Both are in Class 9, C = 1.0.
[0132] T (Expansion Type Weight): "Kezhilian" is a shape similarity expansion, T = 0.9.
[0133] S (Similarity): The glyphs of "Kezhilian" and "Kezhilian" are highly similar. Assume that the calculated S = 0.95.
[0134] St_CN (Strength of Prior Trademark): "Kezhilian" is a general registered trademark with average market popularity. Assume St_CN = 0.2.
[0135] LLM_Success_Coeff: Qwen3-SFT-CN predicts that the inherent registration success coefficient of "Kezhilian" in Class 9 in China is 0.75.
[0136] Bad Words: Not hit.
[0137] r_single (“Kezhilian” vs “Kezhilian”) = 0.2×(1 / 3⋅1.0 + 1 / 3⋅0.9 + 1 / 3⋅0.95) ≈ 0.2×0.95 ≈ 0.19。
[0138] Assume this is the only significant conflict, r_max = 0.19。
[0139] = 0.75×(1 - 0.19) = 0.75×0.81 = 0.6075。
[0140] 5. The result summary and recommendation module summarizes the result: For the target trademark “Kezhilian” (Chinese market), P = 0.6075, and the inherent success coefficient of the LLM is 0.75. The main conflicting trademark is “Kezhilian”, and the conflict details.
[0141] 6. The data visualization and export module visually displays P = 0.6075 (for the Chinese market), lists the conflict information of “Kezhilian”, and allows the export of a PDF report.
[0142] 7. The database and model management module provides trademark data for China, the Qwen3 model, the Qwen3 - SFT - CN model, the Chinese stop - word library, and relevant rules in this process.
[0143] Through the collaborative work and regional adaptation of the above modules, this system can provide users with comprehensive, efficient, and intelligent trademark registration risk assessment services for specific markets (such as China, the United States).
[0144] Embodiment 2 Please refer to Figure 2 and Figure 3 , which shows a multi - dimensional vocabulary expansion query method for trademark registration risk assessment provided by an embodiment of the present invention, including the following steps: Step S010, the user starts a trademark risk assessment task.
[0145] Step S020, the user inputs and configures parameters: The user inputs the trademark name to be queried on the system interface, selects the goods / services category (Nice Classification) to be applied for, and critically selects the target market / country (such as China or the United States). The system receives and validates the above - input information.
[0146] Principle: Ensure that all subsequent analyses are based on the region specified by the user to achieve regionalized and accurate assessment.
[0147] Step S030: Regionalized LLM Selection and Multi-Dimensional Vocabulary Expansion: Based on the user's selected target market, the system dynamically calls the corresponding pre-trained LLM from the database and model management module (e.g., Qwen3 for the Chinese market, Gemma3 for the US market). This LLM, incorporating regional linguistic and cultural characteristics, expands the target trademark across multiple dimensions, including visual similarity, phonetic similarity, semantic similarity, and cross-lingual expansion, generating a vocabulary containing multiple potential close variants.
[0148] Principle: LLM leverages its powerful semantic understanding and generation capabilities, combined with regionalized knowledge, to maximize the simulation of various similar forms of trademarks that may appear in real business scenarios, thereby improving the comprehensiveness of the search.
[0149] Step S040 : Aggregating the target vocabulary input by the user and the extended vocabulary set to form a target extended vocabulary set.
[0150] Step S050, Batch Parallel Search: The system performs a high-speed, parallel search of the expanded vocabulary set generated by the regionalized multidimensional vocabulary expansion engine module and the target vocabulary entered by the user against the official trademark database corresponding to the target market stored in the database and model management module, targeting the user-specified category and related similar categories. The search process is optimized using DFA principles to ensure efficiency.
[0151] The searched keywords are a target expanded vocabulary set, including the target vocabulary input by the user and the expanded vocabulary set generated by the regionalized multi-dimensional vocabulary expansion engine module.
[0152] Principle: Through efficient string matching algorithms and parallel processing mechanisms, prior trademark records that are identical or highly similar to the expanded word can be quickly screened from massive trademark data.
[0153] Step S060: Determine whether a similar trademark has been retrieved. If no prior similar trademark is retrieved, the process may jump to step S080, and the system will give an evaluation conclusion mainly based on the inherent registrability of the target trademark (from LLM_Success_Coeff).
[0154] If a prior similar trademark is retrieved, the process goes to step S070 to perform a detailed conflict risk assessment.
[0155] Step S070, Regionalized Five-Dimensional + Penalty Item Risk Assessment and Registrability Prediction: This is the core processing step of the present invention. For each retrieved prior similar trademark, the system will call the regionalized five-dimensional + penalty item risk assessment and registrability prediction module to perform a detailed risk assessment in combination with the target trademark (or its related extensions). The regionalized five-dimensional + penalty item risk assessment and registrability prediction module will calculate the individual risk coefficient of each potential conflict. , and finally summarize the maximum conflict risk And the comprehensive registration probability P.
[0156] Principle: Through a complex model that is multi-dimensional, quantitative, and regionally adaptable, it simulates the judgment logic of trademark examiners, comprehensively considers various factors that affect trademark similarity and registrability, and provides a scientific risk assessment.
[0157] See also Figure 3 , showing the conflict risk scoring sub-process, which calculates the individual risk coefficient for each pair of "target trademark (or one of its extensions) and retrieved prior trademark" , and finally summarize the maximum conflict risk And the comprehensive registration probability P, the specific steps are as follows: Step S070.1, start risk assessment (for a single portfolio and target market): sub-process starts.
[0158] Step S070.2, input data: The sub-process receives as input the target word, previous word, target market currently being evaluated, product / service category, and expansion type of the target word (such as similar in form, similar in sound, etc.).
[0159] Step S070.3, calculate the Nice Classification compatibility (C): the system queries the "Classification of Similar Goods and Services" or relevant case rules for the current target market stored in the database and model management module, determines the degree of compatibility between the category designated by the target word and the category designated by the prior trademark, and gives a quantitative score C (for example, 1.0 represents the same category, 0.7 represents similarity, etc.).
[0160] Principle: Similarity of goods / services is one of the prerequisites for trademark infringement or confusion.
[0161] Step S070.4, determining the extension type weight (T): Based on the extension type of the target word (labeled by the regionalized multidimensional vocabulary extension engine module), and querying the weight rules for the target market in the database and model management module, a weight value T is assigned to the extension type (for example, the weight of sound similarity may be higher than the weight of semantic similarity).
[0162] Rationale: Different types of approximation have different degrees of likelihood of causing confusion.
[0163] Step S070.5, calculate literal / semantic similarity (S): Utilize a combination of string edit distance (e.g., Levenshtein), phonetic encoding comparison (e.g., Double-Metaphone or Chinese pinyin sequence comparison), and semantic similarity calculated using a regionalized LLM (e.g., cosine similarity) to calculate the literal and semantic similarity between the target word and the preceding word, yielding a score S in the range of 0-1.
[0164] Principle: The similarity of the trademark logo itself is the core of judging similarity.
[0165] Step S070.6, assessing the strength of the prior trademark (St): querying the trademark strength assessment rules and relevant data for the target market stored in the database and model management module (such as the USPTO's Section 15 status, China's well-known trademark records, etc.), comprehensively assessing the prior trademark's distinctiveness, fame, legal protection status, etc., and providing a quantitative strength score St (for example, St_CN or St_US).
[0166] Principle: The stronger the prior trademark, the larger its scope of protection is generally, and the higher the risk of being similar to it.
[0167] Step S070.7, obtain the inherent registration success coefficient (LLM_Success_Coeff): call the LLM (such as Qwen3-SFT-CN or Gemma3-SFT-US) stored in the database and model management module, which has the official review opinion SFT of the target market, input the target trademark name, category and market, predict its own inherent registrability (without considering the conflict with specific prior trademarks), and obtain a probability value LLM_Success_Coeff in the range of 0-1.
[0168] Principle: AI is used to learn from a large number of real review cases to determine whether the trademark itself meets the absolute grounds for registration (such as distinctiveness, non-descriptiveness, etc.).
[0169] Step S070.8, check for banned words (Bad Words): query the banned word list for the target market stored in the database and model management module.
[0170] Step S070.9, determine whether the target word hits the banned word: If so, the final registrable probability P of the target trademark will be assigned an extremely low value (such as 0.01) or directly marked as "prohibited from registration". The sub-process risk assessment of this combination can be terminated early or marked as the highest risk.
[0171] If not, continue with the subsequent calculations.
[0172] ,in , , is the weight of each factor (for example, they can all be set to 1 / 3, or learned from a large amount of case data, with the sum being 1). As a multiplying factor, the greater risk of prior trademarks is emphasized.
[0173] Step S070.10, calculate the single conflict risk coefficient : If no banned words are found, then the formula (for example, ,in , , is the preset or learned weight) to calculate the conflict risk coefficient of the current "target word vs. previous word" combination.
[0174] Principle: Integrate the four dimensions of C, T, S, and St to quantify the approximate risk level between the two.
[0175] Step S070.11, output intermediate results: The sub-process outputs the current combination of Nice Classification compatibility C, extension type weight T, literal / semantic similarity S, prior trademark strength St, and single conflict risk coefficient. Value, as well as the inherent registration success coefficient LLM_Success_Coeff (an inherent value, which remains unchanged for the same target trademark in the same market query) and BadWords hit situation.
[0176] Step S070.12, end single risk assessment: the assessment of this single combination is completed. The regionalized five-dimensional + penalty item risk assessment and registrability prediction module iterates this sub-process until all retrieved prior trademarks are assessed.
[0177] In all single combinations After the calculation is completed, the regionalized five-dimensional + penalty item risk assessment and registrability prediction module will determine the largest After the calculation is completed, the regionalized five-dimensional + penalty item risk assessment and registrability prediction module determines the largest After the calculation is complete, the results are then combined with LLM_Success_Coeff and the stopword check.
[0178] Step 070.13, calculate the conflict risk coefficient r_single and the final registration probability P of the target trademark, and transmit them to the result summary and recommendation module: Among them, the value of the conflict risk coefficient r_single is the maximum conflict risk ; The final registration probability P is value.
[0179] Step S080, Results Summary and Recommendation: After all prior trademarks are evaluated, the results summary and recommendation module is responsible for summarizing the analysis results. If risks exist, high-risk prior trademarks and their conflict details are listed. If there are no significant conflicts, the inherent registrability assessment is primarily presented. Results of no significant conflicts (or only inherent risks) are summarized and a structured risk report is generated.
[0180] Principle: Convert complex analytical data into conclusions and reports that are easy for users to understand.
[0181] Step S090, data visualization and export: the data visualization and export module displays the report content generated by the result summary and recommendation module in a visual form such as a chart on the user interface, and provides a function of exporting the complete report to a format such as PDF.
[0182] Principle: Improve user experience, make risk information more intuitive, and facilitate user decision-making and archiving.
[0183] Step S100, end: the user obtains the evaluation result and the process ends.
[0184] The multi-dimensional vocabulary expansion query system and method for trademark registration risk assessment described above achieves the following beneficial effects: 1. Significantly improve the comprehensiveness of trademark searches and the accuracy of risk identification, effectively reducing the risk of underreporting: Through a regionalized multi-dimensional vocabulary expansion engine, the present invention can intelligently generate a comprehensive set of variants of the target trademark in multiple dimensions, including form, sound, meaning, colloquialisms, internet slang, and cross-language transcriptions, targeting specific target markets (such as China and the United States) by utilizing advanced large language models (such as Qwen3 and Gemma3) and combining them with regionalized language and cultural characteristics.
[0185] For example, for Chinese trademarks, the Qwen3 model can better understand and generate homophones, synonyms, and even internet buzzword variants that conform to Chinese language habits; for English trademarks, the Gemma3 model can effectively handle their spelling variants and slang usage.
[0186] This regionalized and multi-dimensional expansion capability far exceeds traditional methods, and can maximize the coverage of potential similar trademarks and significantly improve the search recall rate.
[0187] Combining the regionalized five-dimensional + penalty item risk assessment with the weight assignment of different extension types (T values) in the registrability prediction module, as well as the precise calculation of literal / semantic similarity (S values), the risk assessment is made more accurate, thereby effectively reducing the risk of underreporting due to incomplete retrieval.
[0188] 2. Significantly improve the efficiency of trademark search and risk assessment, achieving "one-click" rapid response: Automated process: After the user inputs the target trademark and selects the market and category, the system automatically completes the entire process of vocabulary expansion, full-category parallel search, risk calculation and prediction.
[0189] High-speed parallel search: The batch parallel search module uses the DFA principle to optimize the search engine and establishes efficient indexes for the target market and Nice Classification. It can perform high-speed parallel searches on the large number of expanded terms generated by the regionalized multidimensional vocabulary expansion engine module, ensuring that results are returned in a short time.
[0190] Instant risk assessment: The regionalized five-dimensional + penalty risk assessment and registrability prediction module can quickly perform complex five-dimensional risk calculations on each retrieved prior trademark.
[0191] End users can achieve "one-click operation" and obtain a comprehensive risk assessment report within seconds or tens of seconds, which greatly improves efficiency compared to the traditional manual method of several hours.
[0192] 3. Provide quantitative, objective, and regionally adapted risk assessment and registrability prediction to assist users in making accurate decisions: Five-dimensional + penalty risk model: The regionalized five-dimensional + penalty risk assessment and registrability prediction module innovatively constructs a "five-dimensional" evaluation system that includes Nice Classification compatibility C, extended type weight T, literal / semantic similarity S, regionalized prior trademark strength St_CN / St_US, and the registration success coefficient LLM_Success_Coeff predicted by LLM based on the official review opinion SFT, and combines it with a regionalized "bad words" penalty mechanism.
[0193] Quantitative output: The model ultimately outputs a registration probability P value between 0 and 1, as well as the inherent registration success coefficient directly given by the SFT LLM. These quantitative indicators provide users with a clear and intuitive risk reference.
[0194] Regionalized precision assessment: The calculation of trademark strength St fully takes into account the different legal practices of China and the United States (such as Article 15 of the United States and the well-known trademark system of China).
[0195] The prediction of LLM success factor is based on SFT learning of official review opinions of each target market (USPTO / CNIPA), making the prediction more in line with local examination standards.
[0196] The banned word database is also constructed separately according to the laws and regulations of China and the United States.
[0197] This deep regionalization makes the risk assessment results more targeted and practical, avoiding the potential incompatibility of the "one-size-fits-all" model in different jurisdictions.
[0198] Improved decision-making quality: Users no longer simply receive a list of similar trademarks, but instead obtain a specific risk probability based on multi-dimensional data and intelligent analysis, enabling them to make decisions more quickly and confidently on whether to apply, modify, or abandon a trademark.
[0199] 4. Significantly improve user experience, providing intuitive visual reports and convenient export functions: Results summary and recommendations: Clearly summarize key risk information and provide preliminary modification suggestions based on the source of risk.
[0200] Data visualization and export: Risk assessment results are intuitively displayed through various visualization methods, including dashboards, risk lists, and even regionalized risk heat maps. Users can clearly see the final P value, major conflicting trademarks, and their risk composition.
[0201] One-click report generation: Detailed risk assessment reports, including all analysis elements, can be exported to standard formats such as PDF with one click, making them easy to archive, share, and analyze. The target market for the assessment is clearly marked in the report.
[0202] 5. The system has good scalability and maintainability: Modular design: The system adopts a modular architecture with independent functions and clear interfaces for each module, which facilitates future expansion to new target markets (such as the EU and Japan). Only the corresponding regionalized LLM, SFT model, trademark database, banned word library and rule parameters need to be developed or configured.
[0203] Separation of Models and Data: The database and model management module centrally manages various regionalized data and models, facilitating updates, maintenance, and version control, ensuring the accuracy and timeliness of system assessments. For example, this module can regularly update national trademark databases, update LLM models, and adjust risk assessment parameters and banned word lists based on the latest legal precedents.
[0204] This design ensures long-term system availability and continuous optimization capabilities.
[0205] It should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A multi-dimensional vocabulary expansion query system for trademark registration risk assessment, used to build a multi-dimensional, regionalized, quantitative conflict risk assessment and registrability prediction model, characterized by: include: The user input and parameter configuration module is used to receive the user's query request and configuration parameters, and to verify and encapsulate the input data; the configuration parameters include the name of the trademark to be queried, the category of goods and services to be applied for, and the target market area; The regionalized multi-dimensional vocabulary expansion engine module is used to call the corresponding pre-trained large language model (LLM) based on the target market region of the configuration parameters. In combination with the specific language and cultural characteristics of the target market, it automatically generates a set of text variants of the target vocabulary in multiple dimensions such as form, sound, meaning, colloquialism, and cross-language. DFA search module for high-speed, parallel exact string matching searches of all or a user-specified subset of the Nice Classification in the official trademark databases of the target market area; The regionalized five-dimensional plus penalty risk assessment and registrability prediction module is used to conduct a multi-dimensional quantitative conflict risk assessment on each retrieved prior trademark, combined with the user's proposed target trademark and its extension, based on the legal practices and examination standards of the target market. It then uses the supervised fine-tuned SFT LLM to predict the registrability of the target trademark in the target market region. The result summary and recommendation module is used to generate a structured risk report and list the prior similar trademarks that pose the main risk and detailed conflict information. The risk report includes the target trademark, query parameters, the final registration probability P value, and the inherent success coefficient predicted by LLM; Data visualization and export module, used to display risk assessment reports in a friendly graphical interface and provide report export function; The database and model management module is used to store and manage the regionalized data and models required for system operation, and provide data retrieval and caching services.
2. The multi-dimensional vocabulary expansion query system for trademark registration risk assessment according to claim 1, characterized in that: The user input and parameter configuration module includes: A user interface interaction unit is used to receive query requests and configuration parameters input by users through a web interface or an API interface; An input verification unit is used to perform basic verification on the length of the entered trademark name, special characters, selected service category, and target market area; Parameter encapsulation unit, used to encapsulate valid user input and configuration parameters into structured data objects; The task distribution unit is used to transfer the encapsulated data object to the regionalized multi-dimensional vocabulary expansion engine module.
3. The multi-dimensional vocabulary expansion query system for trademark registration risk assessment according to claim 1, characterized in that: The regionalized multidimensional vocabulary expansion engine module includes: An LLM calling and variant generation unit is used to call the selected LLM through the API to obtain a preliminary list of expanded words with regionalized features; The traditional algorithm supplement and regionalized refinement unit is used to generate multi-dimensional vocabulary variants, including form similarity expansion, phonetic similarity expansion, semantic similarity expansion, colloquial expansion, and cross-language expansion, based on the regionalized Prompt algorithm guiding the LLM to conduct multi-dimensional vocabulary expansion around the target trademark; The result deduplication and screening unit is used to screen the variants generated by LLM and supplemented by traditional algorithms based on the initial similarity and relevance with the original word and the language habits of the target market, and to control the scale and quality of the expanded set; The structured output unit is used to construct the final expanded word set into a list of multi-dimensional information including similarity in form, sound, meaning, colloquialisms, and cross-language information, forming an expanded vocabulary set with regional characteristics.
4. The multi-dimensional vocabulary expansion query system for trademark registration risk assessment according to claim 1, characterized in that: The DFA retrieval module includes: A parallel search execution unit is used to achieve high-speed retrieval by using a deterministic finite automaton (DFA) algorithm, state machine precompilation, and single-scan technology. A finite state machine is constructed using a predefined state set Q, a character set Σ, and a state transition function δ. The keyword library is constructed as a Trie tree structure. A single scan can simultaneously detect all predefined keywords, allowing rapid searches for prior trademarks identical to the expanded term within the vast trademark database. The result collection and preliminary filtering unit is used to collect all retrieved prior trademark records and perform preliminary filtering based on application date, trademark status, and target market.
5. The multi-dimensional vocabulary expansion query system for trademark registration risk assessment according to claim 1, characterized in that: The regionalized five-dimensional + penalty item risk assessment and registrability prediction module includes: An iterative evaluation unit is used to conduct a five-dimensional evaluation of each target trademark or combination of a word extension and a prior trademark based on the target market. The five-dimensional evaluation includes Nice Classification compatibility (C), extension type weight (T), literal / semantic similarity (S), prior trademark strength (St), and the inherent registration success coefficient (LLM_Success_Coeff) predicted by the LLM. The banned word penalty mechanism unit is used to evaluate the target trademark based on the built-in banned word database of the target market area. If the core part of the target trademark hits the banned word list of the target market area, the final registration probability P is forced to be set to an extremely low value; The comprehensive conflict risk coefficient and registrable probability calculation unit is used to calculate the conflict risk coefficient r_single and the final registrable probability P of the target trademark based on the evaluation results obtained by the iterative evaluation unit and the banned word penalty mechanism unit.
6. The multi-dimensional vocabulary expansion query system for trademark registration risk assessment according to claim 1, characterized in that: The result summary and recommendation module includes: A data aggregation and sorting unit is used to sort prior conflicting trademarks according to the conflict risk coefficient r_single value from high to low, or the final registration probability P value from low to high; The report generation unit is used to fill in the target trademark information, query parameters, final registration probability P value, inherent registration success coefficient LLM_Success_Coeff, and high-risk trademark list according to the preset template. The high-risk trademark list includes name, application number, category, Nice Classification compatibility C, extended type weight T, literal / semantic similarity S, prior trademark strength St, conflict risk coefficient r_single, and registration country / region.
7. The multi-dimensional vocabulary expansion query system for trademark registration risk assessment according to claim 1, characterized in that: The data visualization and export module includes: A visualization unit, which uses a front-end chart library to visually display the report data generated by the result summary and recommendation module in the form of a dashboard, risk list, or regional risk heat map, and indicates the target market area. If multi-market evaluation is supported, it provides switching or parallel display of evaluation results; The report export unit is used to export the complete risk assessment report including all charts and detailed data into the standard PDF format with one click.
8. The multi-dimensional vocabulary expansion query system for trademark registration risk assessment according to claim 1, characterized in that: The database and model management module includes: Regionalized trademark databases, used to store and regularly update official trademark data for target market regions, including trademark names, images, applications / registrations, applicants, designated goods / services classifications, legal status, application / registration dates, and trademark strength information; Regionalized model library, used to store pre-trained LLM models and regionalized LLM models after SFT; User database, used to store user information, query history, and configuration preferences; A regionalized banned word database, which stores banned word lists, legal basis, and severity levels for each target market region; Regionalized rules and parameter library, used to store the Nice Classification compatibility rules, extended type weights, and trademark strength calculation rules for each target market region.
9. A multidimensional vocabulary expansion query method for trademark registration risk assessment, providing users with a trademark registration risk assessment service for a specific market through a multidimensional vocabulary expansion query system for trademark registration risk assessment according to any one of claims 1 to 8, characterized in that: The method comprises the following steps: Step 1: The user initiates a query request and enters configuration parameters; Step 2: Based on the target market region selected by the user, the corresponding pre-trained LLM is called. Combining the regional language and cultural characteristics, the target trademark is expanded in multiple dimensions, including form, sound, meaning, and across languages, to generate an expanded vocabulary set containing multiple potential close variants. Step 3: Aggregate the target vocabulary input by the user and the extended vocabulary set to form a target extended vocabulary set; Step 4: Perform a high-speed parallel search of the official trademark database corresponding to the target market area for the user-specified category and related similar categories; Step 5: Determine whether a prior similar trademark is found. If no prior similar trademark is found, proceed to Step 8. If a prior similar trademark is found, proceed to Step 6. Step 6: For each similar trademark retrieved, conduct a detailed risk assessment in combination with the target trademark or the expanded vocabulary set, and calculate the individual risk coefficient of each potential conflict. , and finally summarize the maximum conflict risk and the comprehensive registration probability P; Step 7: After all prior trademarks have been evaluated, the analysis results are summarized and a structured risk report is generated. Based on the evaluation results, a risk assessment is conducted to determine whether there is a risk. If there is a risk, the high-risk prior trademarks and their conflict details, as well as the inherent registrability risk, are listed. If there is no significant conflict, only the inherent registrability risk is displayed. Step 8: Generate a structured risk report; Step 9: Data visualization and export; the generated report content is displayed on the user interface in visual forms such as charts, and the function of exporting the complete report to PDF format is provided.
10. The multi-dimensional vocabulary expansion query method for trademark registration risk assessment according to claim 9, characterized in that: Step 6: For each similar trademark retrieved, a detailed risk assessment is conducted in combination with the target trademark or the expanded vocabulary set, and the individual risk coefficient of each potential conflict is calculated. , and finally summarize the maximum conflict risk The steps of calculating the comprehensive registration probability P specifically include: Step 6.1, receiving input parameter values, including target word, previous word, target market area, product / service category, and target word expansion type; Step 6.2: Calculate the Nice Classification compatibility, C. Based on the "Classification of Similar Goods and Services" for the target market area or case law, determine whether the target goods and the goods with the prior trademark belong to the same or similar categories. A comprehensive assessment of similarity is made based on the nature, use, and distribution of the goods / services, and the score is mapped to a similarity score. The value of C ranges from 0 to 1. Step 6.3: Determine the weight T of the expansion type. When determining the impact of each dimension of lexical expansion on the likelihood of confusion, T ranges from 0 to 1. The weight T can be fine-tuned based on regional legal practice, giving different weights to similar-sounding expansions, similar-sounding expansions, regionalized transcriptions, and similar-sense expansions. Step 6.4, calculate the literal / semantic similarity S; calculate the similarity between the target trademark or its extension and the prior trademark name, including: Similar extension: The normalized Levenshtein distance is shown in formula (1): Sortho=1−(LevenshteinDistance / max(len(str1),len(str2)))(1); Among them, Sortho represents the normalized literal similarity, and the value range of Sortho is 0~1; LevenshteinDistance represents the edit distance between two strings, which is the minimum number of single-character edits required to convert one string to another; len(str1) / len(str2) represents the length of the first string and the second string respectively; Phonetic extension: Determines the degree of pronunciation similarity between the target trademark or extension and the prior trademark name. The pronunciation similarity Sphonetic value ranges from 0 to 1. When the regionalized phonetic code is a perfect match, Sphonetic = 1.0; Similarity expansion, which is the semantic similarity between the target word and the original word given when LLM is used for vocabulary expansion, or the semantic similarity between the target word and the prior trademark calculated through regionalized LLM; Step 6.5: Evaluate the strength of the prior trademark (St); determine the strength of the prior trademark's distinctiveness, reputation, and legal protection in the country of its registration; Step 6.6: Obtain the inherent registration success coefficient LLM_Success_Coeff; the large language model (SFT) fine-tuned by the official trademark review opinions in the target market region is used to directly predict the inherent registrability and initial success probability of the target trademark in the corresponding regional market; Step 6.7, check for banned words and build a banned word list for the target market area; Step 6.8, determine whether the target word hits the banned word; if the target trademark hits the banned word in the target market area, then, =0.01, or a more detailed penalty value determined by the severity of the banned words, otherwise, = ; Step 6.9, calculate the single conflict risk coefficient The single conflict risk coefficient is the degree of conflict between the target trademark or expansion word and each prior trademark. The calculation formula is shown in formula (2): (2); in, 、 、 are the Nice Classification compatibility C, the extension type weight T, and the literal / semantic similarity S weights, respectively. 、 、 The sum of the three is 1; is the strength of the prior trademark, The value range is 0~1; Step 6.10, output intermediate results; output the current combination of Nice Classification compatibility C, extended type weight T, literal / semantic similarity S, prior trademark strength St, and single conflict risk coefficient Value, as well as the inherent registration success coefficient LLM_Success_Coeff, and Bad Words hit situation. Among them, the inherent registration success coefficient LLM_Success_Coeff is fixed for the same target trademark in the same target market area query; Step 6.11: End the single risk assessment. Repeat steps 6.2 to 6.10 for each combination of the target trademark or word extension and each prior trademark until all retrieved prior trademarks have been assessed. Step 6.12, Determine the Maximum Conflict Risk and the comprehensive registration probability P; 1) Maximum risk of conflict : is the conflict coefficient of all single lines The maximum value among the above is obtained by comparing the target trademark with all prior trademarks. The maximum value of is calculated as shown in formula (3): = max( , ,..., )(3); If there is no conflict, then =0; 2) The final registration probability P includes: i) Basic probability , the calculation formula is shown in formula (4): =LLM_Success_Coeff(4); Among them, LLM_Success_Coeff is the inherent registration success coefficient predicted by LLM; ii) Conflict-adjusted probability , the calculation formula is shown in formula (5): = ×(1− )(5); Among them, 1− ≥0; if ≥1 then 1− Considered as 0; iii) Banned word penalty: If the target trademark hits the banned words in the target market area, then, =0.01, or a more detailed penalty value determined by the severity of the banned words, otherwise, = ; Step 6.13: Calculate the conflict risk coefficient r_single and the final registration probability P of the target trademark and transmit them to the result summary and recommendation module: Among them, the value of the conflict risk coefficient r_single is the maximum conflict risk ; The final registration probability P is value.
Citation Information
Cited By
Trademark examination method and device, electronic equipment and storage medium
CN121833919A