Trademark batch retrieval method and system based on big data and ai assistance
By using a multimodal trademark analysis model assisted by big data and AI, the problem of insufficient integration of graphic and textual features in trademark retrieval has been solved, enabling efficient and accurate batch trademark retrieval and meeting the retrieval needs of enterprises.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies struggle to effectively integrate graphic and textual features during trademark retrieval, resulting in insufficient retrieval accuracy and low retrieval efficiency, failing to meet enterprises' needs for bulk trademark retrieval.
A multimodal trademark analysis model based on big data and AI is adopted. By acquiring a batch of trademarks to be searched and the trademark search condition text, graphic attribute features and search condition features are extracted, a feature matrix is constructed and matrix operations are performed to calculate the multimodal correlation degree and generate batch search results.
It enables efficient batch retrieval and accurate sorting of trademarks, significantly improving retrieval efficiency and accuracy, and meeting the actual needs of enterprises.
Smart Images

Figure CN121412411B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a method and system for batch trademark retrieval based on big data and AI assistance. Background Technology
[0002] Currently, traditional methods for trademark retrieval rely heavily on manual comparison or simple keyword matching, which is insufficient to meet the demand for efficient retrieval of large numbers of trademarks. Existing technologies often fail to effectively integrate graphic, textual, and other feature elements when processing multimodal trademark information, leading to insufficient retrieval accuracy. Furthermore, faced with a large number of trademarks to be retrieved and complex search conditions, retrieval efficiency is low, and it is difficult to achieve accurate ranking of search results, failing to meet the actual needs of enterprises and other users for trademark strategy and risk assessment. Therefore, how to leverage big data and artificial intelligence technologies to improve the efficiency and accuracy of bulk trademark retrieval has become an urgent technical problem to be solved. Summary of the Invention
[0003] The purpose of this invention is to provide a method and system for batch trademark retrieval based on big data and AI assistance.
[0004] In a first aspect, embodiments of the present invention provide a method for batch trademark retrieval based on big data and AI assistance, comprising: acquiring a batch set of trademarks to be retrieved, the batch set of trademarks to be retrieved including multiple trademark subjects to be retrieved; acquiring multiple trademark retrieval condition texts; inputting the multiple trademark retrieval condition texts into a trained multimodal trademark analysis model to obtain retrieval condition features of the trademark retrieval targets included in each trademark retrieval condition text; inputting each trademark subject to be retrieved in the batch set of trademarks to be retrieved into the multimodal trademark analysis model, the multimodal trademark analysis model processing each trademark subject to be retrieved to obtain corresponding graphic attribute features, forming a set of trademark features to be retrieved; constructing a retrieval condition feature matrix of the multiple trademark retrieval condition texts and a graphic attribute feature matrix of the set of trademark features to be retrieved, and batch calculating the multimodal correlation degree between the retrieval condition features of each trademark retrieval condition text and the graphic attribute features of each trademark subject to be retrieved through matrix operations to obtain a batch correlation degree matrix; determining the trademark retrieval targets of the trademark subjects to be retrieved corresponding to each trademark retrieval condition text according to the batch correlation degree matrix, and generating batch retrieval results, the batch retrieval results including multiple trademark subjects to be retrieved matched by each trademark retrieval condition text and their correlation degree ranking.
[0005] In a second aspect, embodiments of the present invention provide a server system, including a server, the server being used to execute the method described in the first aspect.
[0006] Compared to existing technologies, the beneficial effects provided by this invention include: A trademark batch retrieval method and system based on big data and AI assistance, as disclosed in this invention, includes: acquiring a batch set of trademarks to be retrieved and multiple trademark retrieval condition texts; using a trained multimodal trademark analysis model to extract retrieval condition features from the retrieval condition texts and graphic attribute features of the trademark subjects to be retrieved; constructing a retrieval condition feature matrix and a graphic attribute feature matrix; and calculating multimodal correlation degrees in batches through matrix operations to obtain a batch correlation degree matrix; determining retrieval targets based on the batch correlation degree matrix and generating batch retrieval results including correlation degree ranking. This method, through multimodal feature fusion and matrix operations, achieves efficient batch retrieval and accurate ranking of trademarks, effectively improving retrieval efficiency and accuracy, and meeting practical application needs. Attached Figure Description
[0007] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 A flowchart illustrating the steps of the trademark batch retrieval method based on big data and AI assistance provided in an embodiment of the present invention;
[0009] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0011] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings. In order to solve the technical problems in the aforementioned background art, Figure 1This is a flowchart illustrating the trademark batch retrieval method based on big data and AI assistance provided in this embodiment. The following is a detailed description of this trademark batch retrieval method based on big data and AI assistance. Step S201: Obtain a batch set of trademarks to be retrieved, which includes multiple trademark entities. Step S202: Obtain multiple trademark retrieval condition texts. Step S203: Input the multiple trademark retrieval condition texts into a trained multimodal trademark analysis model to obtain the retrieval condition features of the trademark retrieval target included in each trademark retrieval condition text. Step S204: Input each trademark entity in the batch set of trademarks to be retrieved into the multimodal trademark analysis model. The multimodal trademark analysis model processes each trademark entity to obtain corresponding graphic attribute features, forming the trademark to be retrieved. Feature set; Step S205, construct the retrieval condition feature matrix of the multiple trademark retrieval condition texts and the graphic attribute feature matrix of the trademark feature set to be retrieved, and calculate the multimodal correlation degree between the retrieval condition features of each trademark retrieval condition text and the graphic attribute features of each trademark subject to be retrieved in batches through matrix operations to obtain a batch correlation degree matrix; Step S206, determine the trademark retrieval target of the trademark subject to be retrieved corresponding to each trademark retrieval condition text according to the batch correlation degree matrix, and generate batch retrieval results, the batch retrieval results including multiple trademark subjects to be retrieved matched by each trademark retrieval condition text and their correlation degree ranking.
[0012] In an embodiment of the present invention, for example, the server system of a certain intellectual property service platform needs to provide trademark batch search services for 5 corporate clients (Company A to Company E). Each company submits the trademark subject to be searched and the corresponding search condition text. The server needs to complete the batch search and return the results through a multimodal trademark analysis model.
[0013] The server first obtains a batch of trademarks to be searched through a pre-defined interface (such as a dedicated API for enterprise clients, a web upload port, or a database synchronization channel). This set contains multiple trademark entities to be searched, each of which can be a graphic file (such as a trademark image in PNG or JPG format), official graphic data corresponding to a trademark registration number, or a trademark material combining text and graphics.
[0014] Specifically, Company A submitted three trademark entities for search: ① a purely graphic trademark "Flying Dove" (PNG format, 500×500 resolution, featuring a flying dove outlined in black lines on a white background, with outstretched wings and a streamlined body); ② a text + graphic trademark "Green Sprout Technology LOGO" (containing the green text "Green Sprout" and two sprout graphics below); ③ a purely graphic trademark "Circular Gear Combination" (blue circular background with three intersecting silver gear graphics inside). Company B submitted two trademark entities: ① a graphic trademark "Water Drop Shield" (a blue water drop shape with a gold shield outline embedded within); ② a text trademark "Lightning Energy" (pure text "Lightning Energy," without graphics). Companies C through E each submitted 1-2 trademark entities. After the server aggregated these, a batch of 10 trademark entities for search was formed, and each trademark entity was assigned a unique identifier (e.g., "Search-001" to "Search-010"), which was stored in the "Search Queue" of the distributed database.
[0015] The server retrieves multiple trademark search criteria texts corresponding to the batch of trademarks to be searched through the enterprise client's system interface or preset templates. Each search criteria text is entered by the enterprise client according to their own needs, including information such as search objectives, element restrictions, and exclusion conditions.
[0016] For example, Company A submitted two search criteria texts: Search criterion text 1: "Search for registered trademarks similar to our company's 'Flying Pigeon' (pending search-001) graphic, requiring the inclusion of 'wing elements' and 'streamlined body' elements, excluding purely textual trademarks and trademarks with the color red"; Search criterion text 2: "Search for similar trademarks in Class 9 (scientific instruments) registered as 'Circular Gear Combination' (pending search-003), requiring the graphic to contain 'gear' elements, the color to be blue or silver, and the registration status to be 'valid'." Company B submitted one search criteria text: Search criterion text 3: "Search for trademarks similar to the 'Water Drop Shield' (pending search-004) graphic, requiring the inclusion of 'water drop outline' and 'shield embedded' elements, prioritizing trademarks in Class 35 (advertising and sales).
[0017] Enterprises C through E each submit one search condition text. The server aggregates these, obtaining a total of five trademark search condition texts (numbered "Search Condition-T01" to "Search Condition-T05"), which are stored in the "Search Task Table" and associated with the corresponding enterprise ID and the identifier of the trademark subject to be searched. The server calls the trained multimodal trademark analysis model, inputting the five trademark search condition texts obtained in step 2 into the model sequentially, and outputting the search condition features (including search description features, shape perception features, and search instruction features) corresponding to each text. The core function of the multimodal trademark analysis model is to perform multimodal parsing of text-based search conditions: first, semantic information is extracted through the natural language processing module; then, the core elements in the search target are broken down through the trademark element parsing engine to generate element appearance perception data (describing visual attributes) and search instruction information (element classification identifier + search instruction framework), which are finally fused into a search condition feature vector.
[0018] Taking search condition text 1 ("Search for registered trademarks similar to the graphic of 'Flying Pigeon' (searchable-001), requiring the inclusion of 'wing element' and 'streamlined body' elements, excluding purely text trademarks and trademarks with the color red") as an example, the model processing process is as follows: Trademark element analysis: Three core trademark elements are separated to form a set of trademark elements: Element 1: "Wing element" (graphic element, the core visual element of the search target); Element 2: "Streamlined body" (graphic element, the core visual element of the search target); Element 3: "Exclude purely text trademarks and trademarks with the color red" (instruction element, search restriction condition). Generate appearance perception data: For graphic elements (elements 1 and 2), structured data describing their visual attributes are generated through a pre-trained visual attribute dictionary: Appearance perception data for element 1 "wing element": "Visual attributes: symmetrical structure, simulated feather texture (line density 5-8 lines / cm), spread angle 110°-130°, wingtip curvature radius 0.5-1cm"; Appearance perception data for element 2 "streamlined body": "Visual attributes: ratio of major axis to minor axis 3:1, head is arc-shaped (radius 0.8cm), torso line smoothness ≥0.9 (based on curve fitting error)".
[0019] Generate search instruction information: For all elements, based on the element classification identifier (the preset trademark element classification system, such as "graphic element - animal organ" and "exclusion instruction - type restriction") combined with the search instruction framework (such as "include", "exclude", "prioritize", etc.), the following is generated: The element classification identifier of element 1 is "graphic element - animal organ", and the search instruction framework is "include", so the search instruction information is: "[graphic element - animal organ] includes 'wing element'"; The element classification identifier of element 3 is "exclusion instruction - type restriction" and "exclusion instruction - color restriction", and the search instruction framework is "exclude", so the search instruction information is: "[exclusion instruction - type restriction] excludes pure text trademarks; [exclusion instruction - color restriction] excludes trademarks with the color red".
[0020] Extracting search condition features: Search description features: The overall semantics of search condition text 1 are encoded to obtain a 1024-dimensional vector, which includes semantic information such as "similar to a flying pigeon graphic", "wings + streamlined body", and "exclude pure text / red"; Shape perception features: Based on appearance perception data, a 1024-dimensional vector is generated through a visual feature extraction network (such as a ResNet-50 pre-trained model) to encode visual features such as the outline of the wings and the streamlined curve of the body; Search instruction features: The search instruction information is encoded to obtain a 512-dimensional vector, which includes instruction logic such as "included elements" and "exclusion conditions".
[0021] The server performs the same processing on the remaining four search condition texts (such as search condition texts 2 to 5), and obtains the corresponding search condition features (all of which are fusion vectors of search description features + shape-aware features + search instruction features), which are stored in the "search feature library" and associated with the search condition text numbers.
[0022] The server retrieves each trademark subject from the batch of trademarks to be retrieved in the "queue to be retrieved" of the distributed database, and inputs them sequentially into the trained multimodal trademark analysis model. The model extracts visual features from them to obtain the corresponding graphic attribute features, and finally forms a feature set of trademarks to be retrieved.
[0023] Graphical attribute features are high-dimensional vectors describing the visual essence of a trademark, including core visual attributes such as outline, color, texture, and element layout. Taking the trademark subject "Flying Pigeon" (searchable-001) as an example, the model processing is as follows: Preprocessing: Convert the PNG image to a standardized size (512×512 pixels) and perform grayscale normalization (if it is a color image, retain the RGB channels); Element analysis and visual feature extraction: Locate the core area of the trademark subject through an object detection network (such as Faster R-CNN) (excluding background interference), and then extract the following features through a multi-scale feature fusion module: Outline features: Polygon fitting parameters of the overall outline of the pigeon (12 vertices, maximum side length 3.2cm); Color features: RGB color distribution (90% black, 10% white background, no other colors); Texture features: Line density of the wing area (7 lines / cm), smoothness of the body area (curve fitting error 0.12); Element layout features: Relative position of the wings and body (angle between the wing midline and the body midline is 85°).
[0024] Generate graphic attribute features: The above visual features are fused into a 1024-dimensional vector (matching the dimension of the search condition features) as the graphic attribute features of the target -001.
[0025] The server performs the same processing on the remaining 9 trademark subjects to be searched (such as "Green Sprout Technology LOGO" and "Circular Gear Combination"). For purely graphic trademarks (such as search-003 "Circular Gear Combination"), the focus is on extracting features such as the shape, color (blue, silver), and quantity (3) of the gears. For text + graphic trademarks (such as search-002 "Green Sprout Technology LOGO"), only the visual features of the graphic part (sprout) are extracted, ignoring the text part. For purely text trademarks (such as search-006 "Lightning Energy"), the model marks its "graphic attribute features as empty" (because there are no graphic elements). The server summarizes the graphic attribute features of the 10 trademark subjects to be searched, forming a set of trademark features to be searched, which is stored in the "feature database to be searched" and associated with the trademark subject identifier.
[0026] The server retrieves the retrieval condition features of multiple trademark retrieval condition texts from the "retrieval feature library" to construct a retrieval condition feature matrix; it retrieves the graphic attribute features of the trademark feature set to be retrieved from the "feature library to be retrieved" to construct a graphic attribute feature matrix; and it calculates the multimodal correlation degree between the two in batches through matrix operations to obtain a batch correlation degree matrix.
[0027] Construct the retrieval condition feature matrix: Assuming there are 5 trademark retrieval condition texts (T01-T05), each retrieval condition feature is a 2560-dimensional vector (retrieval description feature 1024-dimensional + shape perception feature 1024-dimensional + retrieval instruction feature 512-dimensional, obtained by weighted fusion), then the retrieval condition feature matrix is a 5×2560 matrix (rows: retrieval condition text number, columns: feature dimension).
[0028] Constructing the graphic attribute feature matrix: The batch of trademarks to be searched contains 10 trademark subjects (to be searched -001 to -010), and each graphic attribute feature is a 1024-dimensional vector (the graphic attribute feature vector of pure text trademarks is all 0). Then the graphic attribute feature matrix is a 10×1024 matrix (rows: trademark subject identifiers to be searched, columns: feature dimensions).
[0029] The server calculates multimodal correlation by combining matrix multiplication with cosine similarity: First, the retrieval condition feature matrix (5×2560) and the transpose of the graphic attribute feature matrix (1024×10) are aligned in dimensions (the graphic attribute feature vector is mapped from 1024 dimensions to 2560 dimensions through a feature mapping layer) to obtain a 5×10 intermediate matrix. Then, the cosine similarity is calculated for each element (the value ranges from 0 to 1, and the higher the value, the stronger the correlation). Finally, a batch correlation matrix (5×10) is obtained, where the matrix element (i,j) represents the multimodal correlation between the i-th retrieval condition text and the j-th trademark subject to be retrieved.
[0030] For example, the results of the correlation calculation between search condition text 1 (T01) and the subject of the trademark to be searched are as follows (partial examples): Search-001 ("Flying Pigeon"): correlation 0.93 (highest self-graphic matching degree); Search-004 ("Water Drop Shield"): correlation 0.21 (no wings / streamlined body elements); Search-006 ("Lightning Energy" pure text trademark): correlation 0 (graphic attribute features are empty, triggering exclusion conditions).
[0031] The server sorts the row vectors corresponding to each search condition text in the batch relevance matrix (i.e., the relevance of the search condition to all the trademark subjects to be searched) in descending order, selects the trademark subjects to be searched with a relevance of ≥0.6 as the matching results (too low a relevance is considered irrelevant), and marks the relevance value and sorting to generate batch search results.
[0032] Taking the search condition text 1 (T01) as an example, its corresponding row vectors, after sorting, are as follows: Search-001 (relevance 0.93); Search-008 (Company D's "White Dove Messenger" graphic trademark, relevance 0.78, containing wings and a streamlined body, color is black and white); Search-003 ("Circular Gear Combination", relevance 0.62, although there is no dove element, some outline lines are similar to the streamlined shape of wings, color is blue (not red), not excluded); Search-007 (Company C's "Flying Bird LOGO", relevance 0.58, below the threshold of 0.6, not included in the results).
[0033] The server generates the search results for search condition text 1: "Matching trademark subjects to be searched: 1. To be searched-001 (relevance 0.93); 2. To be searched-008 (relevance 0.78); 3. To be searched-003 (relevance 0.62)", and attaches a thumbnail of each trademark subject and the basis for the relevance calculation (such as "containing wing elements (match 0.85), streamlined body (match 0.91)").
[0034] After performing the same processing on the remaining four search criteria texts, the server summarizes all results into a batch search results table, which includes fields such as "search criteria text number", "matching trademark subject identifier", "relevance", "sorting", and "element matching details". The table is returned to the enterprise customer through the API interface and stored in the "search results archive", thus completing the batch search process.
[0035] Through the above steps, the server utilizes a multimodal trademark analysis model to achieve cross-modal association between search criteria text and the subject of the trademark to be searched. Combined with matrix operations, it efficiently processes batch tasks, significantly improving the efficiency and accuracy of trademark retrieval and meeting the batch retrieval needs of enterprise clients.
[0036] In this embodiment of the invention, the retrieval condition features of each trademark retrieval target include at least one of the following features: retrieval description features, shape perception features, and retrieval instruction features; the multimodal trademark analysis model is used to: obtain the retrieval description features based on the trademark retrieval condition text, or to perform trademark element parsing on the trademark retrieval condition text to obtain a set of trademark elements of the trademark retrieval condition text, generate appearance perception data and retrieval instruction information for each trademark element in the set of trademark elements, obtain the shape perception features of each trademark element based on the appearance perception data of each trademark element in the set of trademark elements, and obtain the retrieval instruction features of each trademark element based on the retrieval instruction information of each trademark element in the set of trademark elements; each trademark element in the set of trademark elements corresponds to a trademark retrieval target, the appearance perception data is used to describe the visual attributes of the trademark element, and the retrieval instruction information is composed of the element classification identifier of the trademark element combined with the retrieval instruction framework.
[0037] In this embodiment of the invention, for example, when the server processes the trademark search condition text to obtain search condition features, it first calls the natural language understanding module in the multimodal trademark analysis model to perform global semantic encoding on the input trademark search condition text for the extraction of search description features. For example, for the search condition text 1 submitted by Company A, "Search for registered trademarks that are similar in graphic to our company's 'Flying Pigeon' (to be searched-001), requiring the inclusion of 'wing elements' and 'streamlined body' elements, excluding pure text trademarks and trademarks with the color red", the model will first perform word segmentation processing to obtain keyword sequences such as "search", "flying pigeon", "graphic similarity", "include", "wing elements", "streamlined body", "exclude", "pure text trademark", and "color red". Through the attention mechanism, it captures the inter-word dependencies. For example, the co-occurrence weight of "wing elements" and "streamlined body" is 0.82, which is higher than the co-occurrence weight of "exclude" and "color red" (0.65), thereby determining the core search target. These semantic information were then mapped to a 1024-dimensional vector space to generate a retrieval description feature vector. The activation value of the dimension corresponding to "graphic similarity" was 0.91, the activation values of the dimensions corresponding to "wing element" and "streamlined body" were 0.87 and 0.85 respectively, and the activation value of the dimension corresponding to "exclusion" was 0.72, which intuitively reflects the importance of each element in the text.
[0038] Next, the server uses the model's trademark element parsing engine to parse the trademark search condition text, resulting in a set of trademark elements, each corresponding to a trademark search target. For the aforementioned search condition text 1, after parsing, a set containing three trademark elements is formed: the core graphic elements "wing element" and "streamlined body," and the instruction element "exclude pure text trademarks and trademarks with red color." For the graphic elements in the set, the server generates appearance perception data to describe their visual attributes. For example, the appearance perception data for the "wing element" includes "structural type: symmetrical structure, texture features: feather-like lines (density 6 lines / cm), geometric parameters: spread angle 120°, wingtip curvature radius 0.8cm, wing root width 2.3cm," etc. The appearance perception data for the "streamlined body" includes descriptions such as "major axis to minor axis ratio 3:1, head is arc-shaped (radius 0.8cm), torso line smoothness ≥0.9," etc. Then, based on these appearance perception data, the shape perception features of each trademark element are obtained. The model calls the pre-trained visual feature extraction network to convert the above structured appearance perception data into a 1024-dimensional shape perception feature vector. For example, the "spreading angle 120°" of the "wing element" is mapped to the 357th dimension of the vector with an activation value of 0.89, and the "symmetric structure" is mapped to the 512th dimension with an activation value of 0.93.
[0039] Simultaneously, for each trademark element in the trademark element set, the server constructs search instruction information based on its element classification identifier and search instruction framework, and obtains search instruction characteristics accordingly. The element classification identifier is a preset trademark element classification system. For example, the element classification identifier for "wing element" is "graphic element - animal organ," and the element classification identifier for "exclude pure text trademarks" is "exclusion instruction - type restriction." The search instruction framework includes "include" and "exclude," etc. Therefore, the search instruction information for "wing element" is "[graphic element - animal organ] includes 'wing element'," and the search instruction information for "exclude pure text trademarks and trademarks with red color" is "[exclusion instruction - type restriction] exclude pure text trademarks; [exclusion instruction - color restriction] exclude trademarks with red color." The server uses an instruction encoding module to convert these retrieval instruction information into a 512-dimensional retrieval instruction feature vector. The "contains" operation corresponds to dimensions 0-255 with an activation value of 0.92, while the "excludes" operation corresponds to dimensions 256-511 with an activation value of 0.88. Furthermore, the activation strength of the "color restriction" sub-instruction (0.75) is higher than that of the "type restriction" sub-instruction (0.68). In this way, the server ultimately obtains retrieval condition features that include at least one of the following: retrieval description features, shape-aware features, and retrieval instruction features, providing multimodal feature support for subsequent trademark association calculations.
[0040] In this embodiment of the invention, the multimodal trademark analysis model is obtained in the following ways, and can be implemented through the following examples.
[0041] The sample dataset is input into a pre-trained multimodal trademark analysis model to obtain the graphic attribute features, retrieval description features, shape perception features, and retrieval instruction features of at least two trademark samples included in the sample dataset. Each trademark sample includes a trademark subject and trademark representation data of the trademark subject. The multimodal trademark analysis model is used to: parse the trademark representation data to obtain a set of trademark elements of the trademark sample; generate appearance perception data and retrieval instruction information of the trademark elements in the set of trademark elements of the trademark sample; obtain the retrieval description features of the trademark sample based on the trademark representation data of the trademark sample; obtain the shape perception features of the trademark sample based on the appearance perception data of the trademark elements in the set of trademark elements of the trademark sample; obtain the retrieval instruction features of the trademark sample based on the retrieval instruction information of the trademark elements in the set of trademark elements of the trademark sample; and obtain the graphic attribute features based on the trademark subject of the trademark sample. The appearance perception data is used to describe the visual attributes of the trademark elements, and the retrieval instruction information is composed of the element classification identifier of the trademark elements combined with the retrieval instruction framework.
[0042] The multimodal correlation between each pair of trademark samples is determined based on the graphic attribute features, retrieval description features, shape perception features, and retrieval instruction features of the at least two trademark samples; the prior confidence distribution correlation between each pair of trademark samples is obtained; the distribution mapping bias is determined based on the multimodal correlation between each pair of trademark samples and the prior confidence distribution correlation; and the multimodal trademark analysis model is adjusted based on the distribution mapping bias.
[0043] In an embodiment of the present invention, for example, when the server obtains the multimodal trademark analysis model, it first constructs a sample dataset containing 1,000 registered trademark samples. Each trademark sample contains the trademark subject (such as a PNG format graphic file) and trademark representation data (the element description, classification information, etc. officially registered by the Trademark Office). For example, Sample 1 is the trademark "Flying Pigeon" (the main body is a pigeon graphic outlined in black and white lines, with a wing spread angle of 120° and a body major axis to minor axis ratio of 3:1). Its trademark representation data includes "Element Classification: Animals - Birds - Pigeon", "Graphic Elements: Wings (symmetrical structure), Body (streamlined)", and "Registration Category: Class 9 (Scientific Instruments)". Sample 2 is the trademark "White Pigeon Messenger" (the main body is a gray pigeon graphic, with a wing spread angle of 115° and a rounded head). Its representation data includes "Element Classification: Animals - Birds - Pigeon", "Graphic Elements: Wings (semi-spread), Head (rounded)", and "Registration Category: Class 9". Sample 3 is the trademark "Gear Machinery" (the main body is a blue circular background with 3 intersecting silver gears embedded). Its representation data includes "Element Classification: Machinery - Gear", "Graphic Elements: Gear (12 teeth), Circular (5cm in diameter)", and "Registration Category: Class 7 (Mechanical Equipment)". The server inputs the sample dataset into a pre-trained (based on ImageNet and trademark text corpus) multimodal trademark analysis model. The model parses the trademark elements of each sample's trademark representation data. For example, the representation data of sample 1 is parsed to parse the following set of trademark elements: "Graphic element - wings (120° spread angle, symmetrical structure)", "Graphic element - body (streamlined, 3cm major axis / 1cm minor axis)", and "Category element - Class 9". Subsequently, appearance perception data for each element is generated. For example, the appearance perception data for "wings" is "visual attributes: symmetrical structure, line density 6 lines / cm, wingtip curvature radius 0.8cm". Based on this, a 1024-dimensional shape perception feature vector is generated through a visual feature extraction network (ResNet-50) (where "symmetrical structure" corresponds to a dimension activation value of 0.93 and "spreading angle 120°" corresponds to a dimension activation value of 0.89). At the same time, a 1024-dimensional search description feature vector is generated based on the description of "animal-bird-pigeon" in the representation data ("pigeon" corresponds to a dimension activation value of 0.92). A 512-dimensional search instruction feature vector is generated based on the search instruction information "[graphic element-wings] contains" ("contains" operation corresponds to a dimension activation value of 0.85). Visual features are extracted from the trademark main graphic file to generate a 1024-dimensional graphic attribute feature vector (outline feature corresponds to a dimension activation value of 0.91).
[0044] Next, the server calculates the multimodal correlation between pairs of samples based on the graphic attribute features (1024 dimensions), retrieval description features (1024 dimensions), shape-aware features (1024 dimensions), and retrieval instruction features (512 dimensions) of 1000 samples. For example, for samples 1 and 2 (both pigeon images), the server calculates the cosine similarity between the graphic attribute features (1024 dimensions) of sample 1 and the comprehensive semantic features (weighted fusion of retrieval description + shape-aware + retrieval instruction features, 2560 dimensions) of sample 2, resulting in a multimodal correlation of 0.85 (high similarity). For samples 1 and 3 (pigeon vs. gear), since the graphic elements have no overlap, the multimodal correlation is only 0.32 (low similarity). Based on this, the server constructs a 1000×1000 multimodal correlation matrix, where the element at position (sample 1, sample 2) has a value of 0.85, and the element at position (sample 1, sample 3) has a value of 0.32.
[0045] Subsequently, the server obtains the prior confidence distribution correlation between each pair of samples. Taking Sample 1 and Sample 2 as examples, the number of common elements in their trademark element sets is first determined: Sample 1's element set is {wings, body, class 9}, and Sample 2's element set is {wings, head, class 9}. The common elements are "wings" and "class 9", totaling 2. The total number of unique elements is the number of elements in the union of Sample 1 and Sample 2's element sets, i.e., {wings, body, head, class 9}, totaling 4. The mapping ratio between the two is the number of common elements / the total number of unique elements = 2 / 4 = 0.5. The server uses this mapping ratio of 0.5 as the prior confidence distribution correlation between Sample 1 and Sample 2. For Sample 1 and Sample 3, Sample 1's element set {wings, body, class 9} and Sample 3's element set {gear, circle, class 7} have no common elements, so the number of common elements is 0, and the total number of unique elements is 6. The mapping ratio 0 / 6 = 0, therefore the prior confidence distribution correlation is 0. Based on this, the server constructs a 1000×1000 prior confidence distribution matrix, with the element at position (sample 1, sample 2) having a value of 0.5 and the element at position (sample 1, sample 3) having a value of 0.
[0046] The server further determines the distribution mapping bias based on the multimodal correlation matrix and the prior confidence distribution matrix. For example, for sample 1 and sample 2, the absolute difference between the multimodal correlation of 0.85 and the prior confidence distribution correlation of 0.5 is 0.35; for sample 1 and sample 3, the absolute difference between the multimodal correlation of 0.32 and the prior confidence distribution correlation of 0 is 0.32. The server calculates the mean bias of all elements in the matrix, obtaining an initial distribution mapping bias of 0.12. To adjust the model, the server backpropagates this distribution mapping bias to the parameters of each layer of the model: for example, reducing the weight of the dimension corresponding to the "body" element in the shape perception features of sample 1 (from 0.85 to 0.78), and increasing the weight of the dimension corresponding to the "wings" element (from 0.89 to 0.94) to reduce the bias between the multimodal correlation and the prior confidence distribution correlation. After 50 rounds of iterative training, the server recalculates the distribution mapping bias, which drops to 0.05 (preset threshold ≤ 0.08). At this point, the model parameters converge, and the multimodal trademark analysis model training is complete.
[0047] In this embodiment of the invention, obtaining the prior confidence distribution correlation degree between each pair of trademark samples in the at least two trademark samples can be implemented through the following example: Based on the trademark element sets of the at least two trademark samples, determine the trademark element similarity coefficient between each pair of trademark samples, and use the trademark element similarity coefficient as the prior confidence distribution correlation degree between the pairs of trademark samples.
[0048] In an embodiment of the present invention, for example, when the server obtains the prior confidence distribution correlation between at least two trademark samples, it first calculates the similarity coefficient of trademark elements between pairs of samples based on the trademark element set of each trademark sample in the sample dataset. The trademark element set of each trademark sample is generated by the model parsing its trademark representation data, and includes multi-dimensional elements such as graphic elements, category elements, and color elements, and each element has a unique identifier (such as "graphic element - wings - spread angle", "category element - class 9", "color element - no specified color"). For example, the sample dataset contains three typical trademark samples: Sample A is the trademark "Flying Pigeon", whose trademark element set is analyzed as {graphic element - wings (120° spread angle, symmetrical structure), graphic element - body (streamlined, 3cm major axis / 1cm minor axis), category element - Class 9 (scientific instruments), color element - no specified color (only black and white lines)}; Sample B is the trademark "White Pigeon Messenger", whose element set is {graphic element - wings (half spread, 115° angle), graphic element - head (arc shape, radius 0.8cm), category element - Class 9 (scientific instruments), color element - no specified color (grayscale graphic)}; Sample C is the trademark "Gear Machinery", whose element set is {graphic element - gear (12 teeth, 3cm diameter), graphic element - circle (5cm background diameter), category element - Class 7 (mechanical equipment), color element - specified color (blue background, silver gear)}.
[0049] The server performs element comparisons on pairs of samples. First, it determines the "number of common elements" in the trademark element sets of the target sample and the comparison sample, i.e., the number of element identifiers that are completely matched. Taking sample A and sample B as examples, comparing their element sets: Although the "graphic element - wings" have different unfolding angles (120° vs 115°), the element identifier "graphic element - wings" is the same (angle is an attribute parameter and does not affect element identifier matching), so it is determined to be a common element; the identifier for "category element - Class 9" is completely identical, making it a common element; the identifiers for "graphic element - body" in sample A and "graphic element - head" in sample B are different, so they are not common elements; the identifiers for "color element - no specified color" in sample A and "color element - no specified color" in sample B are the same, making them common elements. Therefore, the number of common elements between sample A and sample B is 3 (wings, Class 9, no specified color).
[0050] Next, the server determines the "total number of unique features," which is the number of elements in the union of the feature sets of the target sample and the comparison sample (after deduplication). Sample A has 4 features, and sample B has 4 features. The union of the two contains {wings, body, head, category 9, no specified color}, for a total of 5 unique features. Therefore, the total number of unique features is 5.
[0051] Subsequently, the server calculates the trademark element similarity coefficient: the number of common elements (3) divided by the total number of unique elements (5), resulting in 3 / 5 = 0.6. This coefficient is the prior confidence distribution correlation degree between sample A and sample B.
[0052] For samples A and C, comparing the feature sets: the "graphical element - wings" and "graphical element - body" in sample A are completely different from the "graphical element - gear" and "graphical element - circle" in sample C; the "category element - class 9" in sample A is different from the "category element - class 7" in sample C; the "color element - no specified color" in sample A is different from the "color element - specified color" in sample C. The number of common elements is 0, the total number of unique elements is 8 (4 in sample A + 4 in sample C, with no overlap), and the similarity coefficient is 0 / 8 = 0, meaning the prior confidence distribution correlation is 0.
[0053] The server performs the above calculation on all pairwise samples in the sample dataset (a total of 1000×999 / 2=499500 pairs): For sample pairs containing the same core graphic elements (such as "pigeon wings" and "streamlined body"), the same registration category, and the same color type, the number of common elements is high (usually 3-5), the number of unique elements is low (5-7), and the similarity coefficient is high (0.5-0.8); for sample pairs with completely non-overlapping elements (such as animal graphics vs. mechanical graphics), the similarity coefficient is low (0-0.2). The server stores all the calculated similarity coefficients as a 1000×1000 prior confidence distribution matrix. The element value at the (sample i, sample j) position in the matrix is the prior confidence distribution correlation degree between sample i and sample j, providing benchmark data for subsequent calculation of distribution mapping bias.
[0054] In this embodiment of the invention, determining the trademark element similarity coefficient between two trademark samples based on the trademark element sets of the at least two trademark samples can be implemented through the following example: Determine the number of common elements in the trademark element sets of the target trademark sample and the trademark element sets of the comparison trademark sample; determine the total number of unique elements in the trademark element sets of the target trademark sample and the comparison trademark sample; determine the mapping ratio between the number of common elements and the total number of unique elements in the trademark element sets of the target trademark sample and the comparison trademark sample, and obtain the trademark element similarity coefficient between the target trademark sample and the comparison trademark sample.
[0055] In an embodiment of the present invention, for example, when determining the similarity coefficient of trademark elements between two trademark samples, the server first selects the target trademark sample and the comparison trademark sample from the sample dataset, and performs element comparison and quantitative calculation based on the trademark element sets of the two. Taking the trademarks “Flying Pigeon” (denoted as target sample A) and “White Pigeon Messenger” (denoted as comparison sample B) in the sample dataset as examples, the set of trademark elements of target sample A is parsed as: {graphic element - wings (element identifier: GF-Wing), graphic element - body (element identifier: GF-Body), category element - Class 9 (element identifier: CL-09), color element - no specified color (element identifier: CO-None)}, where each element has a unique identifier (e.g., “GF-Wing” represents “graphic element - wings”, and the identifier is composed of element type + core features, and is unrelated to specific attribute parameters); the set of trademark elements of comparison sample B is: {graphic element - wings (element identifier: GF-Wing), graphic element - head (element identifier: GF-Head), category element - Class 9 (element identifier: CL-09), color element - no specified color (element identifier: CO-None)}.
[0056] The server first determines the "number of common elements," which is the number of elements in the trademark element sets of target sample A and comparison sample B that have completely identical element identifiers. Specifically, the server compares them using preset element identifier matching rules (elements with completely identical identifier strings are considered a match): the "graphic element - wings (GF-Wing)" of target sample A is completely identical to the "graphic element - wings (GF-Wing)" of comparison sample B, and is therefore determined to be a common element; the "category element - Class 9 (CL-09)" of target sample A is identical to the "category element - Class 9 (CL-09)" of comparison sample B, and is therefore determined to be a common element; the "color element - no specified color (CO-None)" of target sample A is identical to the "color element - no specified color (CO-None)" of comparison sample B, and is therefore determined to be a common element; however, the "graphic element - body (GF-Body)" of target sample A and the "graphic element - head (GF-Head)" of comparison sample B have different identifiers ("Body" and "Head" have different core features), and therefore do not constitute a common element. According to statistics, the target sample A and the comparison sample B have three common elements.
[0057] Next, the server determines the "total number of unique elements," which is the total number of unique elements after merging the trademark element sets of target sample A and comparison sample B. Target sample A's element set contains 4 elements, and comparison sample B's element set contains 4 elements. Three of these elements (GF-Wing, CL-09, CO-None) are shared by both, and the remaining elements are "GF-Body," unique to target sample A, and "GF-Head," unique to comparison sample B. Therefore, the merged element set is {GF-Wing, GF-Body, GF-Head, CL-09, CO-None}, with a total of 5 unique elements.
[0058] The server then calculates the "mapping ratio," which is the ratio of the number of common elements to the total number of unique elements. This ratio is the trademark element similarity coefficient between the target sample and the comparison sample. For target sample A and comparison sample B, the number of common elements is 3, the total number of unique elements is 5, and the mapping ratio is 3 / 5 = 0.6. Therefore, the trademark element similarity coefficient between the two is 0.6.
[0059] To further verify, the server selected target sample A and the trademark "Gear Machinery" (denoted as comparison sample C) from the sample dataset for calculation. The element set of target sample A is {GF-Wing, GF-Body, CL-09, CO-None}, and the element set of comparison sample C is {graphic element - gear (element identifier: GF-Gear), graphic element - circle (element identifier: GF-Circle), category element - Class 7 (element identifier: CL-07), color element - specified color (element identifier: CO-Specified)}. After comparing the element identifiers, there are no common elements between the two (GF-Wing≠GF-Gear, CL-09≠CL-07, etc.), and the number of common elements is 0; the merged element set contains 8 unique elements (4+4=8, no overlap), and the total number of unique elements is 8; the mapping ratio is 0 / 8=0, therefore the trademark element similarity coefficient between target sample A and comparison sample C is 0.
[0060] Through the above process, the server calculates the similarity coefficients of trademark elements for all pairs of trademark samples in the sample dataset, providing a quantitative basis for subsequently determining the prior confidence distribution correlation. In this embodiment of the invention, the determination of the distribution mapping bias based on the multimodal correlation and prior confidence distribution correlation between the pairs of trademark samples can be implemented through the following example: Construct a correlation matrix with the number of rows and columns equal to the number of trademark samples based on the multimodal correlation between the pairs of trademark samples; construct a prior confidence distribution matrix with the number of rows and columns equal to the number of trademark samples based on the prior confidence distribution correlation between the pairs of trademark samples; determine the distribution mapping bias based on the prior confidence distribution matrix and the correlation matrix.
[0061] In an embodiment of the invention, for example, when the server determines the distribution mapping bias based on the multimodal correlation degree and prior confidence distribution correlation degree between pairs of trademark samples, it first constructs a correlation degree matrix and a prior confidence distribution matrix based on the trademark samples in the sample dataset, and then quantifies the bias between the two by comparing the matrices. Taking three typical trademark samples in the sample dataset as an example: Sample A is the "Flying Pigeon" trademark (element set: {GF-Wing, GF-Body, CL-09, CO-None}), Sample B is the "White Pigeon Messenger" trademark (element set: {GF-Wing, GF-Head, CL-09, CO-None}), and Sample C is the "Gear Machinery" trademark (element set: {GF-Gear, GF-Circle, CL-07, CO-Specified}), the bias between the pairs of the three is determined by the multimodal correlation degree and prior confidence distribution correlation degree. The multimodal correlation and prior confidence distribution correlation have been calculated through the aforementioned steps: the multimodal correlation between sample A and sample B is 0.85 (high similarity) and the prior confidence distribution correlation is 0.6 (trademark element similarity coefficient); the multimodal correlation between sample A and sample C is 0.32 (low similarity) and the prior confidence distribution correlation is 0 (no common elements); the multimodal correlation between sample B and sample C is 0.28 (low similarity) and the prior confidence distribution correlation is 0 (no common elements); the correlation between each sample and itself is 1 (perfect match).
[0062] The server first constructs a correlation matrix based on multimodal correlation. The number of rows and columns in this matrix equals the number of trademark samples in the sample dataset (here, 3 rows and 3 columns). The element in the i-th row and j-th column represents the multimodal correlation between the i-th sample and the j-th sample. Specifically, for the three samples mentioned above, the correlation matrix M is constructed as follows: .
[0063] Among them, the diagonal elements (such as M[0][0], M[1][1]) are all 1.00, which means that the multimodal correlation between the sample itself and itself is 1 (complete correlation); the off-diagonal elements M[0][1]=0.85 (multimodal correlation between sample A and sample B), M[0][2]=0.32 (multimodal correlation between sample A and sample C), M[1][2]=0.28 (multimodal correlation between sample B and sample C), which is consistent with the previous calculation results.
[0064] Next, the server constructs a prior confidence distribution matrix based on the prior confidence distribution correlation. This matrix has the same number of rows and columns as the number of samples (3 rows and 3 columns), and the element in the i-th row and j-th column represents the prior confidence distribution correlation between the i-th sample and the j-th sample (i.e., the trademark element similarity coefficient). Based on the aforementioned calculations, the prior confidence distribution correlation between sample A and sample B is 0.6, the prior confidence distribution correlation between sample A and sample C, and between sample B and sample C are both 0, and the prior confidence distribution correlation between a sample itself is 1 (the number of common elements in its own element set equals the total number of unique elements, with a mapping ratio of 1). Therefore, the prior confidence distribution matrix P is constructed as follows: .
[0065] Subsequently, the server determines the distribution mapping bias based on the prior confidence distribution matrix P and the correlation matrix M. The bias calculation uses the "mean absolute error" as a quantification metric; that is, it first calculates the absolute difference between corresponding elements in matrix P and matrix M, and then calculates the average of all differences. Specifically:
[0066] The difference between the elements at position (0,0) of matrices P and M is: |1.00 - 1.00| = 0.00; the difference at position (0,1) is: |0.60 - 0.85| = 0.25; the difference at position (0,2) is: |0.00 - 0.32| = 0.32; the difference at position (1,0) is: |0.60 - 0.85| = 0.25 (symmetric to the difference at position (0,1)); the difference at position (1,1) is: |1.00 - 1.00|=0.00; Difference at position (1,2): |0.00-0.28|=0.28; Difference at position (2,0): |0.00-0.32|=0.32 (symmetrical to position (0,2); Difference at position (2,1): |0.00-0.28|=0.28 (symmetrical to position (1,2); Difference at position (2,2): |1.00-1.00|=0.00.
[0067] Sum the differences at the nine positions: 0.00 + 0.25 + 0.32 + 0.25 + 0.00 + 0.28 + 0.32 + 0.28 + 0.00 = 1.70. Divide this sum by the total number of elements, 9, to obtain the mean absolute error, which is approximately 1.70 / 9 ≈ 0.19. The server determines this mean absolute error of 0.19 as the current distribution mapping bias, reflecting the overall deviation between the multimodal correlation matrix and the prior confidence distribution matrix.
[0068] Through the above process, the server completes the calculation of the distribution mapping bias. Subsequently, based on this bias (0.19), the parameters of the multimodal trademark analysis model (such as the weights of the visual feature extraction network, the attention coefficients of the semantic encoding layer, etc.) will be adjusted in reverse to reduce the deviation between the multimodal correlation degree output by the model and the prior knowledge (similarity of trademark elements), thereby improving the accuracy of the model's representation of trademark features.
[0069] In this embodiment of the invention, the correlation matrix includes a first correlation matrix, a second correlation matrix, and a third correlation matrix. The determination of the multimodal correlation between two trademark samples based on the graphic attribute features, retrieval description features, shape-aware features, and retrieval instruction features of the at least two trademark samples can be implemented through the following example.
[0070] The process involves determining a first multimodal correlation degree between graphic attribute features and retrieval description features among all pairs of trademark samples in the at least two trademark samples; determining a second multimodal correlation degree between graphic attribute features and shape-aware features among all pairs of trademark samples in the at least two trademark samples; determining a third multimodal correlation degree between graphic attribute features and retrieval instruction features among all pairs of trademark samples in the at least two trademark samples; and constructing a correlation degree matrix with the number of rows and columns equal to the number of trademark samples based on the multimodal correlation degree between all pairs of trademark samples, including: constructing a first correlation degree matrix with the number of rows and columns equal to the number of trademark samples based on the first multimodal correlation degree between all pairs of trademark samples; constructing a second correlation degree matrix with the number of rows and columns equal to the number of trademark samples based on the second multimodal correlation degree between all pairs of trademark samples; and constructing a third correlation degree matrix with the number of rows and columns equal to the number of trademark samples based on the third multimodal correlation degree between all pairs of trademark samples.
[0071] In this embodiment of the invention, for example, when determining the multimodal correlation degree, the server needs to calculate the correlation degree for different feature combinations of trademark samples and construct the corresponding correlation degree matrix. Taking three trademark samples in the sample dataset (Sample A: "Flying Pigeon", Sample B: "White Pigeon Messenger", Sample C: "Gear Machinery") as an example, each sample has extracted graphic attribute features (1024-dimensional visual feature vector, extracted from the main graphic of the trademark), retrieval description features (1024-dimensional semantic feature vector, extracted from the text description of the trademark representation data), shape perception features (1024-dimensional shape feature vector, extracted from the appearance perception data), and retrieval instruction features (512-dimensional instruction feature vector, aligned to 1024 dimensions through a mapping layer, extracted from the retrieval instruction information). The server calculates the multimodal correlation degree of different feature combinations using cosine similarity and constructs three correlation degree matrices.
[0072] First Relevance Matrix: Based on the Relevance Between Graphic Attribute Features and Search Description Features: The server first calculates the first multimodal correlation between each pair of samples based on "graphic attribute features - search description features". Graphic attribute features reflect the visual essence of the trademark subject (such as outline, color, texture), while search description features reflect the textual semantics of the trademark representation data (such as "pigeon graphic" "Class 9"). The correlation between the two measures the degree of matching between visual and semantic features.
[0073] Taking samples A and B as examples: Sample A's graphic attribute feature vector (denoted as GA) includes visual features such as "symmetrical wing structure" (dimension 512, activation value 0.93) and "streamlined body" (dimension 357, activation value 0.89); Sample B's retrieval description feature vector (denoted as DA) is extracted from its representation data "element classification: animal-bird-pigeon" and "graphic element: wings (semi-spread)," and includes semantic features such as "pigeon" (dimension 289, activation value 0.91) and "wings" (dimension 512, activation value 0.87). The server calculates the cosine similarity between GA and DA, i.e., the first multimodal association degree between sample A and sample B is 0.82. Calculation of Sample A and Sample C: The retrieval description feature vector (DC) of Sample C contains "gear" (dimension 412, activation value 0.95) and "machine" (dimension 678, activation value 0.90), which has a large semantic difference from the GA of Sample A (which contains the visual feature of "pigeon"), and the cosine similarity is 0.35, that is, the first multimodal correlation is 0.35.
[0074] Based on the first multimodal association degree of all pairwise samples, the server constructs a 3×3 first association degree matrix M1 (number of rows and columns = number of samples 3): (The diagonal element is 1.00, which represents a perfect visual and semantic match between the sample itself; M1[0][1]=0.82 is the correlation between sample A and B, and M1[0][2]=0.35 is the correlation between sample A and C).
[0075] Second correlation matrix: Correlation between graphic attribute features and shape-perceived features: The server then calculates the second multimodal correlation between "graphic attribute features and shape-perceived features". Shape-perceived features reflect the shape attributes of trademark elements (such as "wing spread angle" and "body streamline"), and the correlation with graphic attribute features measures the degree of matching of visual shapes.
[0076] The graphic attribute features (GA) of sample A and the shape-aware features (SA) of sample B are as follows: Sample B's SA is extracted from the appearance-aware data "wing spread angle 115°, head arc shape", including "wing angle 115°" (dimension 357, activation value 0.85) and "head arc" (dimension 721, activation value 0.88); Sample A's GA includes "wing angle 120°" (dimension 357, activation value 0.89) and "body streamline" (dimension 486, activation value 0.92). The shape features of the two are highly similar, with a cosine similarity of 0.88, i.e., a second multimodal correlation of 0.88.
[0077] Sample A and Sample C: The shape-aware features (SC) of Sample C include "12 gear teeth" (dimension 543, activation value 0.93) and "circular background" (dimension 215, activation value 0.90). It has no shape intersection with the GA (pigeon shape) of Sample A, and the similarity is 0.29, which is the second multimodal correlation of 0.29.
[0078] Construct the second correlation matrix M2: (M2[0][1]=0.88 is the shape correlation between samples A and B, and M2[0][2]=0.29 is the correlation between samples A and C).
[0079] The third correlation matrix: Based on the correlation between graphic attribute features and retrieval instruction features: Finally, the third multimodal correlation of "graphic attribute features - retrieval instruction features" is calculated. Retrieval instruction features reflect the retrieval restrictions of the trademark (such as "contains wing elements" and "excludes red"), and the correlation measures the degree of matching between visual features and instruction logic. Graphic attribute features (GA) of sample A and retrieval instruction features (IA) of sample B: The IA of sample B is extracted from the representation data "contains wing elements" and "Class 9", containing "contains - wings" (dimension 256, activation value 0.85) and "Class - 09" (dimension 312, activation value 0.90); The GA of sample A conforms to the "wing elements" and "Class 9" instructions, with a cosine similarity of 0.79, that is, a third multimodal correlation of 0.79.
[0080] Sample A and Sample C: The retrieval instruction feature (IC) of Sample C includes "specified color - blue silver" (dimension 401, activation value 0.88) and "category - 07" (dimension 312, activation value 0.92), which conflicts with the GA (no specified color, category 9) instruction of Sample A. The similarity is 0.30, that is, the third multimodal correlation is 0.30.
[0081] Construct the third correlation matrix M3: (M3[0][1]=0.79 is the correlation degree between the instructions of sample A and B, and M3[0][2]=0.30 is the correlation degree between sample A and C).
[0082] Through the above process, the server constructs three correlation matrices corresponding to semantic, shape, and instruction dimensions, respectively, providing a hierarchical comparison basis for subsequent calculation of distribution mapping deviation.
[0083] In this embodiment of the invention, determining the distribution mapping bias based on the prior confidence distribution matrix and the correlation matrix can be implemented through the following example: Determine the category mapping bias between the prior confidence distribution matrix and the first correlation matrix; determine the visual mapping bias between the prior confidence distribution matrix and the second correlation matrix; determine the instruction mapping bias between the prior confidence distribution matrix and the third correlation matrix; determine the aggregate mapping bias by summing the category mapping bias, the visual mapping bias, and the instruction mapping bias; and determine the distribution mapping bias based on the aggregate mapping bias.
[0084] In this embodiment of the invention, for example, when the server determines the distribution mapping bias based on the prior confidence distribution matrix and the correlation matrix, it needs to calculate the bias between the prior confidence distribution matrix and the three correlation matrices (first correlation matrix, second correlation matrix, and third correlation matrix) respectively, and then obtain the final distribution mapping bias through bias aggregation. Taking three trademark samples in the sample dataset (sample A: "Flying Pigeon", sample B: "White Pigeon Messenger", sample C: "Gear Machinery") as an example, the prior confidence distribution matrix P has been constructed (number of rows and columns = 3, element values are the similarity coefficients of trademark elements of each pair of samples), and the three correlation matrices are: the first correlation matrix M1 (graphic attribute feature - retrieval description feature correlation), the second correlation matrix M2 (graphic attribute feature - shape perception feature correlation), and the third correlation matrix M3 (graphic attribute feature - retrieval instruction feature correlation), as shown in the following specific matrices:
[0085] Prior confidence distribution matrix P, three correlation matrices: First correlation matrix M1 (graphical description of correlation): Second correlation matrix M2 (graphic-shape correlation): The third correlation matrix M3 (graphics-instruction correlation): Determine the category mapping bias (the bias between P and M1).
[0086] Category mapping bias measures the deviation between the "graphic-description" association (M1) and the prior feature similarity (P). It is calculated as the average of the absolute differences between corresponding elements in matrix P and M1. The server iterates through all elements in the matrix (3×3=9 positions in total):
[0087] Position (0,0): |1.00-1.00|=0.00; Position (0,1): |0.60-0.82|=0.22; Position (0,2): |0.00-0.35|=0.35; Position (1,0): |0.60-0.82|=0.22 (symmetric to (0,1); Position (1,1): |1.00-1.0 0|=0.00; Position (1,2): |0.00-0.33|=0.33; Position (2,0): |0.00-0.35|=0.35 (symmetric to (0,2)); Position (2,1): |0.00-0.33|=0.33 (symmetric to (1,2)); Position (2,2): |1.00-1.00|=0.00. The sum of all differences is: 0.00+0.22+0.35+0.22+0.00+0.33+0.35+0.33+0.00=2.00, and the average is 2.00 / 9≈0.22, meaning the class mapping bias is 0.22.
[0088] Determine visual mapping bias (the deviation between P and M2): Visual mapping bias measures the degree of deviation between the “figure-shape” correlation (M2) and the prior feature similarity (P), and the calculation method is consistent with that of category mapping bias. The server calculates the absolute difference between the elements of matrices P and M2: at position (0,1): |0.60-0.88|=0.28 (the shape correlation between samples A and B is 0.88, which is higher than the prior correlation of 0.60, so the difference is 0.28); at position (0,2): |0.00-0.29|=0.29 (the shape correlation between samples A and C is 0.29, which is higher than the prior correlation of 0, so the difference is 0.29); the difference at other positions is calculated similarly to the category mapping bias: (0,0)=0.00, (1,0)=0.28, (1,1)=0.00, (1,2)=0.27, (2,0)=0.29, (2,1)=0.27, (2,2)=0.00. The sum of all differences is: 0.00+0.28+0.29+0.28+0.00+0.27+0.29+0.27+0.00=2.18. The average value is 2.18 / 9≈0.24, which means the visual mapping deviation is 0.24.
[0089] Determine the instruction mapping bias (the bias between P and M3): The instruction mapping bias measures the degree of deviation between the “graphic-instruction” correlation (M3) and the prior feature similarity (P), and also uses the mean absolute error. The element differences between matrices P and M3 are calculated as follows: At position (0,1): |0.60-0.79|=0.19 (the instruction correlation degree between samples A and B is 0.79, which is higher than the prior correlation degree of 0.60, and the difference is 0.19); At position (0,2): |0.00-0.30|=0.30 (the instruction correlation degree between samples A and C is 0.30, which is higher than the prior correlation degree of 0, and the difference is 0.30); Differences at other positions: (0,0)=0.00, (1,0)=0.19, (1,1)=0.00, (1,2)=0.28, (2,0)=0.30, (2,1)=0.28, (2,2)=0.00. The sum of all differences is: 0.00+0.19+0.30+0.19+0.00+0.28+0.30+0.28+0.00=1.94. The average value is 1.94 / 9≈0.22, which means the instruction mapping deviation is 0.22.
[0090] Determine the aggregation mapping bias (distribution mapping bias): The server sums the category mapping bias (0.22), visual mapping bias (0.24), and instruction mapping bias (0.22) to obtain the aggregation mapping bias: 0.22 + 0.24 + 0.22 = 0.68. This aggregation mapping bias comprehensively reflects the overall deviation between the multimodal correlation (three dimensions) and the prior element similarity. The server uses this as the distribution mapping bias to subsequently adjust the parameters of the multimodal trademark analysis model (such as reducing the weight of the "wing angle" dimension in shape-aware features and increasing the weight of the "category element" dimension in retrieval description features) to reduce the deviation between the model output and prior knowledge.
[0091] In this embodiment of the invention, the deviation between the prior confidence distribution matrix and the correlation matrix is determined as follows: The cross-entropy of the correlation data corresponding to the same target trademark sample in the prior confidence distribution matrix and the correlation matrix is determined, wherein the correlation matrix is the first correlation matrix, the second correlation matrix, or the third correlation matrix; the sum of the cross-entropies of the correlation data corresponding to all target trademark samples in the prior confidence distribution matrix and the correlation matrix is determined to obtain the first cross-entropy; the cross-entropy of the correlation data corresponding to the same comparison trademark sample in the prior confidence distribution matrix and the correlation matrix is determined; the sum of the cross-entropies of the correlation data corresponding to all comparison trademark samples in the prior confidence distribution matrix and the correlation matrix is determined to obtain the second cross-entropy; the mean of the first cross-entropy and the second cross-entropy is determined to obtain the deviation between the prior confidence distribution matrix and the correlation matrix.
[0092] In this embodiment of the invention, for example, when determining the deviation between the prior confidence distribution matrix and the correlation matrix, the server uses cross-entropy as a quantitative indicator, which is achieved by calculating the bidirectional cross-entropy between the target sample and the comparison sample and taking the average. Taking three trademark samples in the sample dataset (sample A: "Flying Pigeon", sample B: "White Pigeon Messenger", sample C: "Gear Machinery") as an example, the deviation calculation process between the prior confidence distribution matrix P and the first correlation matrix M1 (graphic attribute feature - retrieval description feature correlation) is as follows: Cross-entropy requires the input to be a probability distribution (the sum of the elements is 1), so the row and column data of the prior confidence distribution matrix P and the correlation matrix M1 need to be normalized first. Taking matrix rows as an example (from the perspective of the target sample): Row normalization of the prior confidence distribution matrix P: The first row of P (the prior distribution of the target sample A) is [1.00, 0.60, 0.00], and the row sum is 1.00 + 0.60 + 0.00 = 1.60. After normalization, each element is divided by 1.60 to obtain the prior probability distribution p. A =[1.00 / 1.60=0.625,0.60 / 1.60=0.375,0.00 / 1.60=0]; Row normalization of the correlation matrix M1: The first row of M1 (the model predicted correlation of target sample A) is [1.00,0.82,0.35], and the row sum is 1.00+0.82+0.35=2.17. After normalization, each element is divided by 2.17 to obtain the model predicted distribution q. A =[1.00 / 2.17≈0.46,0.82 / 2.17≈0.38,0.35 / 2.17≈0.16].
[0093] The above normalization process is performed on all rows (target samples) and columns (comparison samples) of matrices P and M1 to ensure that the sum of the elements in each row / column is 1, which satisfies the requirements for cross-entropy calculation.
[0094] The first cross-entropy is the sum of the cross-entropies of the associated data corresponding to all target samples. Target samples refer to the row numbers in the matrix; each row corresponds to one target sample, and its associated data consists of the prior distribution (rows of P) and the model prediction distribution (rows of M1) for that row. The formula for calculating cross-entropy is: .in, Let P be the normalized row vector (prior distribution). This is the normalized row vector of M1 (model prediction distribution). This represents the number of samples (n=3 here).
[0095] Cross-entropy of target sample A (first row): prior distribution p A =[0.625,0.375,0]; Model predicts distribution q A =[0.46,0.38,0.16]; Cross-entropy H(p A,q A H(p) = -[0.625×log(0.46)+0.375×log(0.38)+0×log(0.16)]; Substituting into the logarithm (taking the natural logarithm as an example): log(0.46)≈-0.776, log(0.38)≈-0.967, log(0.16)≈-1.833; A ,q A =-[0.625×(-0.776)+0.375×(-0.967)+0]≈-[-0.485+(-0.363)]=-(-0.848)=0.848; Cross-entropy of target sample B (second row): Second row of P: [0.60,1.00,0.00], after normalization p B =[0.60 / 1.60=0.375,1.00 / 1.60=0.625,0]; The second row of M1: [0.82,1.00,0.33], the row sum = 0.82+1.00+0.33=2.15, after normalization q B =[0.82 / 2.15≈0.381, 1.00 / 2.15≈0.465, 0.33 / 2.15≈0.154]; Cross-entropy H(p B ,q B )=-[0.375×log(0.381)+0.625×log(0.465)+0×log(0.154)]; log(0.381)≈-0.964, log(0.465)≈-0.763; H(p B ,q B =-[0.375×(-0.964)+0.625×(-0.763)]≈-[-0.362+(-0.477)]=-(-0.839)=0.839; Cross-entropy of target sample C (third row): Third row of P: [0.00,0.00,1.00], after normalization p C =[0,0,1.00]; The third row of M1: [0.35,0.33,1.00], the row sum = 0.35+0.33+1.00=1.68, after normalization q C =[0.35 / 1.68≈0.208, 0.33 / 1.68≈0.196, 1.00 / 1.68≈0.596]; Cross-entropy H(p C ,q C )=-[0×log(0.208)+0×log(0.196)+1.00×log(0.596)]; log(0.596)≈-0.518; H(p C ,q C=-[0+0+1.00×(-0.518)]=0.518.
[0096] First cross-entropy sum: First cross-entropy = H(p) A ,q A )+H(p B ,q B )+H(p C ,q C = 0.848 + 0.839 + 0.518 ≈ 2.205; Determine the second cross-entropy (the sum of cross-entropies from the perspective of the comparison samples): The second cross-entropy is the sum of the cross-entropies of the associated data corresponding to all comparison samples. The comparison samples refer to the column indices in the matrix, with each column corresponding to one comparison sample. The associated data is the prior distribution (column of P) and the model prediction distribution (column of M1) of that column. The calculation method is similar to that of the first cross-entropy.
[0097] Compare the cross-entropy of sample A (first column): P's first column: [1.00, 0.60, 0.00], after normalization p colA =[0.625,0.375,0] (consistent with the row distribution of target sample A); the first column of M1: [1.00,0.82,0.35], column sum = 1.00 + 0.82 + 0.35 = 2.17, after normalization q colA =[0.46,0.38,0.16] (consistent with the row prediction distribution of target sample A); cross-entropy H(p colA ,q colA )=H(p A ,q A =0.848. Comparing the cross-entropy of sample B (second column): the second column of P: [0.60, 1.00, 0.00], after normalization, p... colB =[0.375,0.625,0] (consistent with the row distribution of target sample B); the second column of M1: [0.82,1.00,0.33], column sum = 0.82+1.00+0.33=2.15, after normalization q colB =[0.381,0.465,0.154] (consistent with the row prediction distribution of target sample B); cross-entropy H(p colB ,q colB )=H(p B ,q B =0.839.
[0098] Compare the cross-entropy of sample C (third column): P's third column: [0.00, 0.00, 1.00], after normalization p colC=[0,0,1.00] (consistent with the row distribution of the target sample C); the third column of M1: [0.35,0.33,1.00], column sum = 0.35+0.33+1.00=1.68, after normalization q colC =[0.208,0.196,0.596] (consistent with the row prediction distribution of the target sample C); cross-entropy H(p colC ,q colC )=H(p C ,q C =0.518. Total second cross-entropy:
[0099] Second cross-entropy = H(p) colA ,q colA )+H(p colB ,q colB )+H(p colC ,q colC =0.848+0.839+0.518≈2.205.
[0100] The deviation (the mean of the first and second cross-entropies) is determined as follows: Deviation = (First cross-entropy + Second cross-entropy) / 2 = (2.205 + 2.205) / 2 = 2.205. Through the above process, the server calculates that the deviation between the prior confidence distribution matrix P and the first correlation matrix M1 is 2.205. Similarly, the above cross-entropy calculation is performed on the second correlation matrix M2 and the third correlation matrix M3 respectively to obtain the visual mapping deviation and the instruction mapping deviation, which are finally aggregated into the distribution mapping deviation to adjust the model.
[0101] In this embodiment of the invention, the determination of the multimodal correlation between two trademark samples based on the graphic attribute features, retrieval description features, shape-aware features, and retrieval instruction features of the at least two trademark samples can be implemented through the following example.
[0102] The comprehensive semantic features of each trademark sample are determined based on the retrieval description features, shape-aware features, and retrieval instruction features of each trademark sample; the multimodal correlation degree of graphic attribute features and comprehensive semantic features between each pair of trademark samples is determined; the determination of distribution mapping bias based on the multimodal correlation degree and prior confidence distribution between each pair of trademark samples includes: constructing a correlation degree matrix with the number of rows and columns equal to the number of trademark samples based on the multimodal correlation degree between each pair of trademark samples; determining the prior confidence distribution correlation degree between each pair of trademark samples based on the prior confidence distribution between each pair of trademark samples; constructing a prior confidence distribution matrix with the number of rows and columns equal to the number of trademark samples based on the prior confidence distribution correlation degree between each pair of trademark samples; and determining the distribution mapping bias based on the prior confidence distribution matrix and the correlation degree matrix.
[0103] In an embodiment of the present invention, for example, when determining the multimodal correlation between two trademark samples, the server first needs to fuse the semantic class features of each sample to generate comprehensive semantic features, then calculate the correlation degree by matching the graphic attribute features with the comprehensive semantic features, and construct a matrix to determine the distribution mapping bias. Taking three trademark samples in the sample dataset (Sample A: "Flying Pigeon", Sample B: "White Pigeon Messenger", Sample C: "Gear Machinery") as an example, each sample has extracted retrieval description features (1024-dimensional semantic vector, denoted as D), shape perception features (1024-dimensional shape vector, denoted as S), retrieval instruction features (instruction vector aligned to 1024 dimensions by the mapping layer, denoted as I), and graphic attribute features (1024-dimensional visual vector, denoted as G, extracted from the main graphic of the trademark).
[0104] The server assigns weights to the importance of retrieval description features, shape-aware features, and retrieval instruction features, and generates comprehensive semantic features through weighted summation. The weight allocation is based on the following criteria: retrieval description features (0.4, reflecting the core semantics of the text), shape-aware features (0.4, reflecting visual shape details), and retrieval instruction features (0.2, reflecting retrieval restrictions), with the sum of the weights of the three being 1.
[0105] Taking sample A as an example: Retrieve descriptive feature D A Extracted from the representation data "Element Classification: Animal-Bird-Pigeon", the activation value of the dimension corresponding to "Pigeon" in the vector is 0.92, and the activation value of the dimension corresponding to "Category 9" is 0.88; shape-aware feature S A Extracted from appearance perception data "wing spread angle 120°, streamlined body", the vector shows an activation value of 0.93 for the dimension corresponding to "symmetrical structure" and 0.89 for the dimension corresponding to "streamlined". Retrieval instruction feature I A Extracted from the search query "contains wing elements, category 9", and aligned to 1024 dimensions via a mapping layer, the activation value for the dimension corresponding to "contains - wings" is 0.85, and the activation value for the dimension corresponding to "category - 09" is 0.90. Comprehensive semantic feature C A The calculation formula is: [C A =0.4×D A +0.4×S A +0.2×I A ]
[0106] Substituting specific vector values (taking the dimension corresponding to "pigeon" as an example): D A This dimension has a value of 0.92, S AThe value of this dimension is 0.85 (related to the "pigeon body" dimension in shape perception), and the value of IA is 0.80 (related to the "animal element" dimension in the instruction). Therefore, the value of CA is 0.4×0.92+0.4×0.85+0.2×0.80=0.368+0.34+0.16=0.868. Performing the same calculation on each dimension of the 1024-dimensional vector yields the comprehensive semantic feature CA (1024 dimensions) of sample A.
[0107] The calculation process for the comprehensive semantic feature CB of sample B (“White Dove Messenger”) is similar: the activation value of the “dove” dimension in DB is 0.90, the activation value of the “half-spread wings” dimension in SB is 0.88, and the activation value of the “contains wing elements” dimension in IB is 0.83. After weighting, the “dove” related dimension value of CB is 0.4×0.90+0.4×0.88+0.2×0.83=0.36+0.352+0.166=0.878. In the comprehensive semantic feature CC of sample C (“Gear Mechanism”), the “gear” related dimension value is 0.4×0.95 (“gear” semantic in DC)+0.4×0.93 (“12 gear teeth” shape in SC)+0.2×0.92 (“contains gear elements” instruction in IC)=0.38+0.372+0.184=0.936.
[0108] The server calculates the multimodal correlation between pairs of samples using cosine similarity, which is the cosine similarity between the graphic attribute feature vector (G) and the comprehensive semantic feature vector (C), and measures the degree of matching between visual features and fused semantic features.
[0109] Taking samples A and B as examples: Sample A's graphic attribute feature GA includes "symmetrical wing structure" (dimension 512, activation value 0.93) and "streamlined body" (dimension 357, activation value 0.89); Sample B's comprehensive semantic feature CB includes "pigeon" semantics (dimension 289, activation value 0.878) and "half-spread wings" shape (dimension 512, activation value 0.88). Calculate the cosine similarity between GA and CB, and substitute it into the vector dot product: GA·CB = 0.93 × 0.88 (dimension 512) + 0.89 × 0.878 (dimension 289) + ... (sum of products of other dimensions) ≈ 0.82 + 0.783 + ... ≈ 856 (assuming the vector magnitude is 10 for simplified calculation), then the similarity ≈ 856 / (10 × 10) = 0.856, that is, the multimodal correlation between sample A and sample B is 0.86 (rounded to two decimal places).
[0110] The correlation between sample A and sample C is calculated as follows: the similarity between GA (pigeon visual features) and CC (gear semantic features) is very low, with a cosine similarity of ≈0.31; the correlation between sample B and sample C is ≈0.29; the correlation between the sample itself is 1.00 (the image and its own comprehensive semantics are perfectly matched).
[0111] Based on this, construct a correlation matrix M (3 rows and 3 columns, with elements representing the pairwise correlation between samples): Determination of the prior confidence distribution matrix and the deviation of the distribution mapping: The prior confidence distribution matrix P follows the previously constructed trademark element similarity coefficient matrix (number of rows and columns = 3): .
[0112] The server calculates the distribution mapping deviation using the mean absolute error: first, it calculates the absolute difference between corresponding elements in matrices P and M, and then takes the average value. The specific calculations are as follows: (0,0) position: |1.00-1.00|=0.00; (0,1) position: |0.60-0.86|=0.26; (0,2) position: |0.00-0.31|=0.31; (1,0) position: |0.60-0.86|=0.26; (1,1) position: |1.00-1.00|=0.00; (1,2) position: |0.00-0.29|=0.29; (2,0) position: |0.00-0.31|=0.31; (2,1) position: |0.00-0.29|=0.29; (2,2) position: |1.00-1.00|=0.00. The sum of all differences is: 0.00 + 0.26 + 0.31 + 0.26 + 0.00 + 0.29 + 0.31 + 0.29 + 0.00 = 1.72. The average value is 1.72 / 9 ≈ 0.19, meaning the distribution mapping bias is 0.19. The server backpropagates this bias to the model, adjusting the weights of the comprehensive semantic features (e.g., reducing the weight of shape-aware features to 0.35 and increasing the weight of retrieval command features to 0.25) to reduce the deviation between the model output and the similarity of prior elements.
[0113] In this embodiment of the invention, the determination of the comprehensive semantic features of each trademark sample based on the retrieval description features, shape-aware features, and retrieval instruction features of each trademark sample can be implemented through the following examples.
[0114] The comprehensive semantic features of each trademark sample are obtained by linearly superimposing the retrieval description features, shape perception features, and retrieval instruction features of each trademark sample.
[0115] In an embodiment of the present invention, for example, when determining the comprehensive semantic features of each trademark sample, the server uses a weighted average method to integrate the retrieval description features, shape perception features, and retrieval instruction features. The weight allocation is based on the importance of each feature to trademark retrieval: retrieval description features (0.4, reflecting the core semantics of trademark representation data), shape perception features (0.4, reflecting the visual shape details of trademark elements), and retrieval instruction features (0.2, reflecting the instruction logic of retrieval restrictions). The sum of the weights of the three is 1.
[0116] Taking sample A (the trademark "Flying Pigeon") as an example, its retrieval description feature is a 1024-dimensional vector D. A (Extracted from the representation data "Feature Classification: Animal-Bird-Pigeon"), where the activation value of the dimension corresponding to "Pigeon" is 0.92; the shape-aware feature is a 1024-dimensional vector S. A (Extracted from appearance perception data "wing spread angle 120°, streamlined body"), the activation value of the corresponding dimension of the "pigeon" is 0.85; the retrieval instruction feature is a vector I aligned to 1024 dimensions by a mapping layer. A (Extracted from the search query information "contains wing elements, category 9"), the activation value of the corresponding dimension for the "pigeon" is 0.80. The server performs a weighted summation on the corresponding dimensions of these three feature vectors: The comprehensive semantic feature vector C A The dimension value corresponding to the "pigeon" is 0.4 × D. A This dimension value is (0.92) + 0.4 × S A This dimension value is (0.85) + 0.2 × I A The value of this dimension (0.80) = 0.4 × 0.92 + 0.4 × 0.85 + 0.2 × 0.80 = 0.368 + 0.34 + 0.16 = 0.868. Performing the above calculation on all dimensions of the 1024-dimensional vector generates the comprehensive semantic feature vector C of sample A. A .
[0117] The processing procedure for sample B ("White Dove Messenger" trademark) is the same: retrieve descriptive feature D. B The "pigeon" corresponds to a dimension activation value of 0.90, and the shape-aware feature S B The activation value for this dimension is 0.88, and the retrieval instruction feature is I. B The activation value for this dimension is 0.83, and the weighted calculation yields the comprehensive semantic feature C. B The dimension value corresponding to "pigeon" is 0.4 × 0.90 + 0.4 × 0.88 + 0.2 × 0.83 = 0.36 + 0.352 + 0.166 = 0.878. The comprehensive semantic feature C of sample C (the trademark "gear machinery") is... C Based on the dimensions related to "gears", D is calculated. C The activation value for this dimension is 0.95, SC The activation value for this dimension is 0.93, I C The activation value for this dimension is 0.92, and after weighting, the result is 0.4 × 0.95 + 0.4 × 0.93 + 0.2 × 0.92 = 0.38 + 0.372 + 0.184 = 0.936. Through the above weighted averaging process, the server generates a 1024-dimensional comprehensive semantic feature vector for each trademark sample. This vector integrates semantic, shape, and instruction information, providing a unified semantic representation for subsequent multimodal correlation calculations.
[0118] In this embodiment of the invention, before adjusting the multimodal trademark analysis model according to the distribution mapping deviation, the following implementation method is also provided: Determining the correlation degree between the rights confirmation benchmarks of each pair of trademark samples based on the rights confirmation benchmark annotations of the at least two trademark samples; determining a reference deviation based on the multimodal correlation degree between the pairs of trademark samples and the rights confirmation benchmark correlation degree; adjusting the multimodal trademark analysis model according to the distribution mapping deviation includes: determining a global optimization objective of the multimodal trademark analysis model based on the distribution mapping deviation and the reference deviation; adjusting the multimodal trademark analysis model according to the global optimization objective.
[0119] In this embodiment of the invention, for example, before adjusting the multimodal trademark analysis model based on the distribution mapping deviation, the server needs to introduce the trademark office's confirmation results as an objective benchmark, and optimize the model by the deviation (reference deviation) between the correlation degree of the confirmation benchmark and the multimodal correlation degree. Taking three trademark samples in the sample dataset as an example: Sample A is the "Flying Pigeon" trademark (the main graphic element is a pigeon outlined in black and white lines, with a 120° wing spread angle and a streamlined body); Sample B is the "White Pigeon Messenger" trademark (the main graphic element is a gray pigeon with partially spread wings and a rounded head); and Sample C is the "Gear Machinery" trademark (the main graphic element is a blue circular background with three silver gears embedded within). The Trademark Office marks the confirmation benchmarks for these three samples as follows: Sample A and Sample B are judged as "similar trademarks" because their core graphic elements (pigeon body and wing structure) are highly similar; Sample A and Sample C, and Sample B and Sample C are judged as "non-similar trademarks" because their element categories (animal graphics vs. mechanical graphics) are significantly different.
[0120] The server assigns a correlation degree to each pair of samples based on the trademark rights confirmation criteria: if two samples are determined to be "similar trademarks", the correlation degree is recorded as 1 (indicating strong correlation); if they are "non-similar trademarks", it is recorded as 0 (indicating no correlation). Specifically, the correlation degree between the trademark rights confirmation criteria of sample A and sample B is 1, while the correlation degree between the trademark rights confirmation criteria of sample A and sample C, and between sample B and sample C, is 0.
[0121] Reference bias is used to measure the degree of deviation between the model's predicted multimodal correlation and the actual weighting results. The server first obtains the previously calculated pairwise multimodal correlation of samples: the multimodal correlation between sample A and sample B is 0.86 (model-predicted similarity), between sample A and sample C is 0.31, and between sample B and sample C is 0.29. Next, the server calculates the reference deviation by "summing the squared deviations and then averaging": Sample A and Sample B: the correlation between the confirmation benchmark is 1, the model prediction correlation is 0.86, the deviation between the two is 1-0.86=0.14, and the squared deviation is 0.14×0.14=0.0196; Sample A and Sample C: the correlation between the confirmation benchmark is 0, the model prediction correlation is 0.31, the deviation is 0-0.31=-0.31, and the squared deviation is 0.31×0.31=0.0961; Sample B and Sample C: the correlation between the confirmation benchmark is 0, the model prediction correlation is 0.29, the deviation is 0-0.29=-0.29, and the squared deviation is 0.29×0.29=0.0841.
[0122] Add the squares of these three deviations together (0.0196+0.0961+0.0841=0.1998), and then divide by the number of sample pairs (3) to get the reference deviation as 0.1998÷3≈0.067.
[0123] The server combines the distribution mapping bias (previously calculated as 0.19, reflecting the deviation between model predictions and prior element similarity) with the reference bias (0.067, reflecting the deviation between model predictions and actual rights confirmation results) to determine the global optimization objective. Considering that the rights confirmation result is the authoritative judgment of the Trademark Office, its weight is higher than that of prior element similarity (empirical judgment based on element logic). The server assigns 60% weight to the reference bias and 40% weight to the distribution mapping bias. The weighted sum of the two is the global optimization objective (the objective is to minimize this value).
[0124] In the specific calculation, the global optimization objective = 0.6 × reference deviation + 0.4 × distribution mapping deviation = 0.6 × 0.067 + 0.4 × 0.19 ≈ 0.0402 + 0.076 = 0.1162.
[0125] To reduce the global optimization objective, the server adjusted the model parameters: enhanced the weight of core graphic elements: the multimodal correlation between sample A and sample B (0.86) was lower than the weighting benchmark (1), and the sensitivity of the "pigeon body" element needed to be increased. The server increased the weight of the "wing outline" dimension in the shape perception feature from 0.4 to 0.45, while reducing the weight of the "wing angle detail" dimension (from 0.3 to 0.25), so that the model pays more attention to the overall outline rather than the local parameter differences, and the correlation between sample A and sample B increased to 0.92; strengthened the distinguishability of category elements: the correlation between sample A and sample C (0.31) was higher than the weighting benchmark (0), because the two belong to different commodity categories (sample A is Class 9 scientific instruments, sample C is Class 7 mechanical equipment). The server increased the weight of the "product category" dimension in the retrieval description features from 0.35 to 0.4, making the model pay more attention to category differences, and the correlation between sample A and sample C decreased to 0.25; the command filtering threshold was optimized: the filtering threshold of "exclude cross-category trademarks" in the retrieval command features was reduced from 0.8 to 0.75, which enhanced the filtering effect on cross-category samples, and the correlation between sample B and sample C further decreased to 0.27.
[0126] After adjustment, the reference biases are updated as follows: sample AB bias squared (1-0.92)²=0.0064, sample AC bias squared (0-0.25)²=0.0625, sample BC bias squared (0-0.27)²=0.0729, and the average of the three is (0.0064+0.0625+0.0729)÷3≈0.0473; the distribution mapping bias is reduced to 0.15. At this point, the global optimization objective = 0.6×0.0473+0.4×0.15≈0.0284+0.06=0.0884 (significantly reduced from the initial 0.1162), the model parameters converge, and the optimization is complete.
[0127] By introducing the correlation degree of the confirmation benchmark and the reference deviation, the server makes the model adjustment both fit the logic of the elements and conform to the real confirmation standard, thus improving the practical accuracy of the multimodal trademark analysis model.
[0128] In this embodiment of the invention, the determination of the reference deviation based on the multimodal correlation and the correlation of the rights confirmation benchmark between the pairs of trademark samples can be implemented through the following example: A correlation matrix with the number of rows and columns equal to the number of trademark samples is constructed based on the multimodal correlation between the pairs of trademark samples; a benchmark indication matrix with the number of rows and columns equal to the number of trademark samples is constructed based on the correlation of the rights confirmation benchmark between the pairs of trademark samples, wherein the rights confirmation benchmark correlation at the same position as the row number and column number in the benchmark indication matrix is marked as a correlation benchmark, and the remaining rights confirmation benchmark correlations are marked as difference benchmarks; the reference deviation is determined based on the correlation matrix and the benchmark indication matrix.
[0129] In this embodiment of the invention, for example, when determining the reference deviation, the server constructs a correlation matrix and a benchmark indication matrix, and compares the differences between the model's prediction results and the authoritative trademark confirmation benchmark. Taking three trademark samples (Sample A: "Flying Pigeon", Sample B: "White Pigeon Messenger", Sample C: "Gear Machinery") as an example, Sample A and B are determined by the Trademark Office to be similar trademarks (relationship degree of 1 in trademark confirmation benchmark), while A and C, and B and C are not similar trademarks (relationship degree of 0 in trademark confirmation benchmark). The model has calculated the multimodal correlation degree between each pair of samples: AB is 0.86, AC is 0.31, BC is 0.29, and the correlation degree of the sample itself is 1.00.
[0130] The correlation matrix records the multimodal correlations predicted by the model. Each row and column corresponds to three samples, and each element represents the correlation between any two samples. The matrix is as follows: The diagonal element (1.00) represents the correlation between samples themselves, while the off-diagonal element represents the correlation between different samples (e.g., 0.86 in the first row and second column indicates the correlation between samples A and B). The benchmark indicator matrix records the trademark office's confirmation results. Each row and column corresponds to three samples and includes two types of annotations: Correlation benchmark: The diagonal position corresponding to the sample itself, labeled as 1 (indicating absolute correlation); Difference benchmark: The off-diagonal positions corresponding to different samples, labeled as the correlation of the confirmation benchmark (approximately 1, non-approximately 0). The matrix is as follows: In this table, 1 in the first row and second column indicates that samples A and B are similar trademarks, and 0 in the first row and third column indicates that samples A and C are not similar trademarks.
[0131] The reference deviation is calculated by comparing the "difference reference position" (non-diagonal element) of the correlation matrix and the reference indicator matrix. The steps are as follows: Extract the values of the difference reference positions: Sample A and B: correlation matrix value 0.86, reference indicator matrix value 1; Sample A and C: correlation matrix value 0.31, reference indicator matrix value 0; Sample B and C: correlation matrix value 0.29, reference indicator matrix value 0. Calculate the squared deviation for each position: Sample A and B: deviation = 0.86 - 1 = -0.14, squared deviation = (-0.14) × (-0.14) = 0.0196; Sample A and C: deviation = 0.31 - 0 = 0.31, squared deviation = 0.31 × 0.31 = 0.0961; Sample B and C: deviation = 0.29 - 0 = 0.29, squared deviation = 0.29 × 0.29 = 0.0841. Calculate the average of the squared deviations: Add the three squared deviations together (0.0196+0.0961+0.0841=0.1998), divide by the number of difference reference positions (3), and get the reference deviation = 0.1998÷3≈0.067.
[0132] The reference bias of 0.067 reflects the overall deviation between the model's predicted correlation and the authoritative confirmation result from the Trademark Office. The smaller the bias, the closer the model's prediction is to the actual confirmation standard. The server will subsequently combine this bias with the distribution mapping bias to determine the model's global optimization objective and adjust the model parameters to reduce the bias.
[0133] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 performs the aforementioned trademark batch retrieval method based on big data and AI assistance. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0134] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.
Claims
1. A trademark batch retrieval method based on big data and AI assistance, characterized in that, The method comprises: obtaining a batch of trademark sets to be searched, the batch of trademark sets to be searched comprising a plurality of trademark subjects to be searched; obtaining a plurality of trademark search condition texts; inputting the plurality of trademark search condition texts into a trained multi-modal trademark analysis model to obtain search condition features of trademark search targets included in each trademark search condition text; inputting each trademark subject to be searched in the batch of trademark sets to be searched into the multi-modal trademark analysis model, and processing each trademark subject to be searched by the multi-modal trademark analysis model to obtain corresponding graphic attribute features, thereby forming a trademark feature set to be searched; constructing a search condition feature matrix of the plurality of trademark search condition texts and a graphic attribute feature matrix of the trademark feature set to be searched, and batch calculating multi-modal correlation degrees between the search condition features of each trademark search condition text and the graphic attribute features of each trademark subject to be searched through matrix operation, thereby obtaining a batch correlation degree matrix; determining trademark search targets of the trademark subjects to be searched corresponding to each trademark search condition text according to the batch correlation degree matrix, and generating a batch search result, the batch search result comprising a plurality of trademark subjects to be searched matched by each trademark search condition text and an associated degree ranking.
2. The method of claim 1, wherein, The search condition features of each trademark search target comprise at least one of the following features: search description features, shape perception features, and search instruction features; The multi-modal trademark analysis model is configured to: obtain the search description features according to the trademark search condition texts, or perform trademark element analysis on the trademark search condition texts to obtain a trademark element set of the trademark search condition texts, generate appearance perception data and search instruction information of each trademark element in the trademark element set, obtain shape perception features of each trademark element according to the appearance perception data of each trademark element in the trademark element set, and obtain search instruction features of each trademark element according to the search instruction information of each trademark element in the trademark element set; Each trademark element in the trademark element set corresponds to a trademark search target, the appearance perception data is used to describe visual attributes of the trademark element, and the search instruction information is constituted by an element classification identifier of the trademark element combined with a search instruction framework.
3. The method of claim 1, wherein, The multi-modal trademark analysis model is obtained by the following method, comprising: inputting a sample data set into a pre-trained multi-modal trademark analysis model to obtain graphic attribute features, search description features, shape perception features, and search instruction features of at least two trademark samples included in the sample data set, each trademark sample including a trademark subject and trademark representation data of the trademark subject, the multi-modal trademark analysis model being configured to: perform trademark element analysis on the trademark representation data to obtain a trademark element set of the trademark sample, generate appearance perception data and search instruction information of a trademark element in the trademark element set of the trademark sample, acquire search description features of the trademark sample according to the trademark representation data of the trademark sample, acquire shape perception features of the trademark sample according to the appearance perception data of the trademark element in the trademark element set of the trademark sample, acquire search instruction features of the trademark sample according to the search instruction information of the trademark element in the trademark element set of the trademark sample, and acquire graphic attribute features according to the trademark subject of the trademark sample, wherein the appearance perception data is used to describe visual attributes of the trademark element, and the search instruction information is composed of element classification identification of the trademark element and a search instruction framework; determining a multi-modal correlation degree between each pair of trademark samples according to the graphic attribute features, search description features, shape perception features, and search instruction features of the at least two trademark samples; acquiring a prior confidence distribution correlation degree between each pair of trademark samples in the at least two trademark samples; determining a distribution mapping deviation according to the multi-modal correlation degree and the prior confidence distribution correlation degree between each pair of trademark samples, and adjusting the multi-modal trademark analysis model according to the distribution mapping deviation.
4. The method of claim 3, wherein, The acquiring of the prior confidence distribution correlation degree between each pair of trademark samples in the at least two trademark samples includes: determining a common element quantity of trademark element quantities of a trademark element set of a target trademark sample and a trademark element set of a comparison trademark sample; determining a total unique element quantity of the trademark element quantities of the trademark element set of the target trademark sample and the trademark element set of the comparison trademark sample; determining a mapping ratio of the common element quantity and the total unique element quantity of the trademark element quantities of the trademark element set of the target trademark sample and the trademark element set of the comparison trademark sample to obtain a trademark element similarity coefficient of the target trademark sample and the comparison trademark sample, and taking the trademark element similarity coefficient as the prior confidence distribution correlation degree between each pair of trademark samples.
5. The method of claim 3, wherein, The determining of the distribution mapping deviation according to the multi-modal correlation degree and the prior confidence distribution correlation degree between each pair of trademark samples includes: constructing a correlation degree matrix with a row number and a column number both being the number of the trademark samples according to the multi-modal correlation degree between each pair of trademark samples; constructing a prior confidence distribution matrix with a row number and a column number both being the number of the trademark samples according to the prior confidence distribution correlation degree between each pair of trademark samples; determining the distribution mapping deviation according to the prior confidence distribution matrix and the correlation degree matrix.
6. The method of claim 5, wherein, The correlation matrix comprises: a first correlation matrix, a second correlation matrix and a third correlation matrix, the multi-modal correlation between each two trademark samples is determined according to the graphic attribute features, the retrieval description features, the shape perception features and the retrieval instruction features of the at least two trademark samples, comprising: determining the first multi-modal correlation between the graphic attribute features and the retrieval description features of each two trademark samples in the at least two trademark samples; determining the second multi-modal correlation between the graphic attribute features and the shape perception features of each two trademark samples in the at least two trademark samples; determining the third multi-modal correlation between the graphic attribute features and the retrieval instruction features of each two trademark samples in the at least two trademark samples; the correlation matrix with the row number and the column number being the trademark sample number is constructed according to the multi-modal correlation between each two trademark samples, comprising: a first correlation matrix with the row number and the column number being the trademark sample number is constructed according to the first multi-modal correlation between each two trademark samples; a second correlation matrix with the row number and the column number being the trademark sample number is constructed according to the second multi-modal correlation between each two trademark samples; a third correlation matrix with the row number and the column number being the trademark sample number is constructed according to the third multi-modal correlation between each two trademark samples.
7. The method of claim 6, wherein, the distribution mapping deviation is determined according to the prior confidence distribution matrix and the correlation matrix, comprising: determining the category mapping deviation of the prior confidence distribution matrix and the first correlation matrix; determining the visual mapping deviation of the prior confidence distribution matrix and the second correlation matrix; determining the instruction mapping deviation of the prior confidence distribution matrix and the third correlation matrix; the sum of the category mapping deviation, the visual mapping deviation and the instruction mapping deviation is the aggregation mapping deviation, and the distribution mapping deviation is determined according to the aggregation mapping deviation; the method further comprises: the deviation of the prior confidence distribution matrix and the correlation matrix is determined by the following way: determining the cross entropy of the correlation data corresponding to the same target trademark sample of the prior confidence distribution matrix and the correlation matrix, wherein the correlation matrix is the first correlation matrix, the second correlation matrix or the third correlation matrix; determining the sum of the cross entropy of the correlation data corresponding to all target trademark samples of the prior confidence distribution matrix and the correlation matrix, obtaining a first cross entropy; determining the cross entropy of the correlation data corresponding to the same comparative trademark sample of the prior confidence distribution matrix and the correlation matrix; determining the sum of the cross entropy of the correlation data corresponding to all comparative trademark samples of the prior confidence distribution matrix and the correlation matrix, obtaining a second cross entropy; determining the mean value of the first cross entropy and the second cross entropy, obtaining the deviation of the prior confidence distribution matrix and the correlation matrix.
8. The method of claim 3, wherein, the multi-modal correlation between each two trademark samples is determined according to the graphic attribute features, the retrieval description features, the shape perception features and the retrieval instruction features of the at least two trademark samples, comprising: linearly superimposing the retrieval description features, the shape perception features and the retrieval instruction features of each of the trademark samples to obtain a comprehensive semantic feature of each of the trademark samples; determining a multi-modal correlation degree between the graphic attribute features and the comprehensive semantic features of each pair of trademark samples in the at least two trademark samples; the determining of the distribution mapping bias according to the multi-modal correlation degree between each pair of trademark samples and the prior confidence distribution comprises: constructing a correlation degree matrix with the number of rows and columns being the number of trademark samples according to the multi-modal correlation degree between each pair of trademark samples; determining a prior confidence distribution correlation degree between each pair of trademark samples according to the prior confidence distribution between each pair of trademark samples, and constructing a prior confidence distribution matrix with the number of rows and columns being the number of trademark samples according to the prior confidence distribution correlation degree between each pair of trademark samples; determining the distribution mapping bias according to the prior confidence distribution matrix and the correlation degree matrix.
9. The method of claim 3, wherein, According to the distribution mapping bias, the method further comprises: determining a registration benchmark correlation degree between each pair of trademark samples according to the registration benchmark annotations of the at least two trademark samples; constructing a correlation degree matrix with the number of rows and columns being the number of trademark samples according to the multi-modal correlation degree between each pair of trademark samples; constructing a benchmark indication matrix with the number of rows and columns being the number of trademark samples according to the registration benchmark correlation degree between each pair of trademark samples, wherein a registration benchmark correlation degree in the same row and column position in the benchmark indication matrix is annotated as a correlation benchmark, and the rest of the registration benchmark correlation degrees are annotated as difference benchmarks; determining a reference bias according to the correlation degree matrix and the benchmark indication matrix; the adjusting of the multi-modal trademark analysis model according to the distribution mapping bias comprises: determining a global optimization target of the multi-modal trademark analysis model according to the distribution mapping bias and the reference bias; adjusting the multi-modal trademark analysis model according to the global optimization target.
10. A server system, characterized by The server is configured to perform the method of any one of claims 1-9. The server is configured to perform the method of any one of claims 1-9.
Citation Information
Patent Citations
Multi-modal trademark retrieval method and system based on comparative learning algorithm
CN116662599A
Trademark image retrieval method and system based on deep learning
CN120144809A