Multi-modal feature-based luxury identification database construction method and system

By constructing a luxury goods identification database with multimodal characteristics, the problem of insufficient single modal information is solved, comprehensive identification and anti-counterfeiting verification of luxury goods is achieved, and the identification efficiency and accuracy are improved.

CN120067400APending Publication Date: 2025-05-30TECH CENT OF GUANGZHOU CUSTOMS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510126594.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has the problem of insufficient single modal information in the identification of luxury goods, which is difficult to provide a comprehensive basis for identification and requires huge database support.

Method used

The luxury goods identification database construction method based on multimodal features is adopted. By obtaining the official information of the luxury goods collection, multiple unit modal information are extracted, and the pre-constructed encoding method is used for index code encoding, and the product-unit modal path is constructed, authorized information, official information and anti-counterfeiting images are integrated to form a unit modal database.

Benefits of technology

A multi-faceted identification of luxury goods has been achieved, and a detailed database containing official information, sales channels and forged characteristics has been constructed, which has improved the identification efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067400A_ABST
    Figure CN120067400A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal databases, and discloses a multi-modal feature-based luxury identification database construction method and system, and the method comprises the steps: extracting official information based on a luxury, carrying out the index code coding of the official information, obtaining product identification data, sequentially extracting unit modal information from the official information, and carrying out the index code coding of the unit modal information; identifying a target modal key based on the extracted unit modal information, taking the product identification data as a father node and the target modal key as a child node, constructing a product-unit modal path, summarizing the product-unit modal path, obtaining an official information structure tree, constructing an authorization information structure tree and an anti-counterfeiting image structure tree, and obtaining an anti-counterfeiting result; and integrating the authorization information structure tree, the official information structure tree and the anti-counterfeiting image structure tree to obtain unit modal databases corresponding to the luxury, and summarizing the unit modal databases to obtain a luxury identification database based on multi-modal features. According to the method, a luxury official information and pseudo luxury database can be constructed according to the multi-modal data features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal databases, and particularly to a method and system for constructing a luxury goods authentication database based on multimodal features. Background Art

[0002] Luxury goods are sought after for their high selling value and collection value, and are consumer goods with unique, scarce, and rare characteristics. Therefore, counterfeits of luxury goods emerge in an endless stream. Especially with the development of technology, counterfeits are exquisitely made, and accurate identification of materials, processes, and designs is required to authenticate counterfeits.

[0003] Existing luxury goods authentication methods include manual authentication and vision-based authentication technologies. Manual authentication is costly and inefficient, while vision-based authentication technologies can perform authenticity detection in a non-invasive, low-cost, and high-efficiency manner.

[0004] Although the above methods can authenticate luxury goods, there is still a problem that the information of a single modality is not sufficient to provide a comprehensive authentication basis. In particular, whenever a person or software authenticates a luxury good, a large database is required as a support. Only with genuine luxury goods, common counterfeit luxury goods, etc. can a luxury good be authenticated in multiple aspects. Therefore, a database storing official information of luxury goods and counterfeit luxury goods is needed. Summary of the Invention

[0005] The present invention provides a method for constructing a luxury goods authentication database based on multimodal features, and a computer can construct a database of official information of luxury goods and counterfeit luxury goods according to multimodal data features.

[0006] To achieve the above object, a method for constructing a luxury goods authentication database based on multimodal features provided by the present invention includes:

[0007] Obtain a set of luxury goods, and perform the following operations on each luxury good in the set of luxury goods:

[0008] Extract official information based on the luxury good, where the official information includes multiple unit modality information, and the multiple unit information includes: brand, item number, origin, luxury good name, initial production date, release batch, limited quantity information, and single-piece style;

[0009] Use a pre-constructed coding method to encode the official information into index code to obtain product identification data;

[0010] Sequentially extract unit modality information from the official information, identify the target modality key based on the extracted unit modality information, use the product identification data as the parent node, and the target modality key as the child node to construct a product-unit modality path, and summarize the product-unit modality paths to obtain an official information structure tree;

[0011] Construct an authorized information structure tree and an anti-counterfeiting image structure tree based on a pre-built authorized data set and a pre-built anti-counterfeiting image set;

[0012] Integrate the authorized information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree to obtain a unit modal database corresponding to the luxury product;

[0013] Summarize the unit modal database to obtain a luxury product authentication database based on multi-modal features.

[0014] Optionally, the extraction of official information based on the luxury product includes:

[0015] Obtain multiple single-modal storage definitions based on multiple unit modal information, where the unit modal information corresponds one-to-one with the single-modal storage definition, and the single-modal storage definition includes: field name, field comment, field type, and field constraint conditions;

[0016] Extract the single-modal storage definitions from the multiple single-modal storage definitions in sequence, extract the field names based on the extracted single-modal storage definitions, construct initial modal key-value pairs with the field names as keys and the pre-built blank values and field comments as values, where the keys in the initial modal key-value pairs are the field names, the values in the initial modal key-value pairs are the field comments and the blank values, and the data type of the blank value is the field type;

[0017] Summarize the initial modal key-value pairs to obtain an initial modal key-value pair set, confirm the initial modal key-value pair set as the initial official information extraction model of the luxury product, and extract official information based on the initial official information extraction model.

[0018] Optionally, the extraction of official information based on the initial official information extraction model includes:

[0019] Receive the original product data, parse the original product data to obtain product structured data and product unstructured data, where the product structured data includes multiple unit structure data;

[0020] Extract the unit structure data from the product structured data in sequence, identify the structure primary key and the structure value in the unit structure data, identify the primary key character segment corresponding to the structure primary key, and use the pre-built semantic similarity algorithm to calculate the character similarity between the field name corresponding to each initial modal key-value pair in the initial modal key-value pair set and the primary key character segment to obtain a character similarity set;

[0021] Extract the character similarities greater than the preset similarity value from the character similarity set to obtain a high similarity character set. If the high similarity character set is an empty set, then eliminate the unit structure data;

[0022] Otherwise, extract the maximum character similarity from the high similarity character set, convert the structure value using the field type to obtain a filled value corresponding to the field type, and map the filled value to the blank value of the initial modal key-value pair corresponding to the maximum character similarity to obtain a filled modal key-value pair;

[0023] Aggregate the filled modal key-value pairs to obtain multiple filled modal key-value pairs, and use the multiple filled modal key-value pairs to update the initial official information extraction model to obtain a sub-optimal official information extraction model;

[0024] Use the product unstructured data to update the sub-optimal official information extraction model to obtain an optimized official information extraction model, and extract official information based on the optimized official information extraction model.

[0025] Optionally, the using the product unstructured data to update the sub-optimal official information extraction model to obtain an optimized official information extraction model, and extracting official information based on the optimized official information extraction model includes:

[0026] Extract initial text data and image data including text from the product unstructured data, identify the text in the image data including text to obtain image text, and aggregate the image text and the initial text data to obtain text to be analyzed;

[0027] Perform keyword classification based on entity extraction, relationship extraction, and attribute extraction on the text to be analyzed based on the sub-optimal official information extraction model to obtain a keyword group set, where the keyword group set includes multiple keyword groups, and each keyword group includes a main word text and one or more content word texts;

[0028] Convert the keyword groups in the keyword group set into key-value pairs to obtain multiple keyword group key-value pairs, aggregate the multiple keyword group key-value pairs to obtain transformed structured data, and use the transformed structured data to update the sub-optimal official information extraction model to obtain an optimized official information extraction model, where the optimized official information extraction model includes multiple optimized modal key-value pairs, and the keyword group key-value pairs are key-value pairs formed with the main word text as the key and the one or more content word texts as the value;

[0029] Perform data cleaning and integrity verification on the optimized official information extraction model, and after confirming that the optimized official information extraction model after performing data cleaning and integrity verification is a preset complete model, obtain a complete official information model, and extract official information based on the complete official information model.

[0030] Optionally, the performing data cleaning and integrity verification on the optimized official information extraction model, and after confirming that the optimized official information extraction model after performing data cleaning and integrity verification is a preset complete model, obtaining a complete official information model includes:

[0031] Extract the optimized modality key-value pairs from the multiple optimized modality key-value pairs corresponding to the optimized official information extraction model in sequence, and perform the following operations on the extracted optimized modality key-value pairs:

[0032] Identify the values of the optimized modality key-value pairs to obtain an optimized modality value set. Among them, the optimized modality value set includes one or more optimized modality values. When there is one optimized modality value in the optimized modality value set, the optimized modality value is a field annotation. When there are multiple optimized modality values in the optimized modality value set, the multiple optimized modality values include the field annotation, and also include the values in the filling modality key-value pairs corresponding to the optimized modality key-value pairs, or the values in the keyword group key-value pairs corresponding to the optimized modality key-value pairs;

[0033] Remove the optimized modality values corresponding to the field annotations from the optimized modality value set, and summarize the remaining optimized modality values to obtain a to-be-verified modality value set. Determine whether the to-be-verified modality value set meets the field constraint conditions corresponding to the optimized modality key-value pairs. Among them, the optimized modality key-value pairs and the initial modality key-value pairs are in one-to-one correspondence. The field constraint conditions corresponding to the optimized modality key-value pairs are the same as the field constraint conditions of the corresponding initial modality key-value pairs, and the character constraint conditions of all initial modality key-value pairs are: unique value and not null;

[0034] If the to-be-verified modality value set corresponding to each optimized modality key value in the multiple optimized modality key values meets the field constraint conditions, confirm that the optimized official information extraction model is the preset complete model to obtain a complete official information model;

[0035] If the to-be-verified modality value set corresponding to the extracted optimized modality key-value pairs does not meet the field constraint conditions and the to-be-verified modality value set is an empty set, receive the target original data based on the optimized modality key-value pairs, parse the target original data, use the parsed target original data to fill the to-be-verified modality value set to obtain a filled modality value set, and use the filled modality value set to update the optimized modality key-value pairs. Use the updated optimized modality key-value pairs as the optimized modality key-value pairs, and return to the step of extracting the optimized modality key-value pairs from the multiple optimized modality key-value pairs corresponding to the optimized official information extraction model in sequence;

[0036] If the to-be-verified modality value set does not meet the field constraint conditions and the to-be-verified modality value set is not an empty set, confirm that the to-be-verified modality value set includes multiple to-be-verified modality values, integrate the to-be-verified modality value set to obtain a unique modality value, and use the unique modality value to update the optimized modality key-value pairs to obtain the final modality key-value pairs;

[0037] Summarize the final modality key-value pairs to obtain a complete official information model.

[0038] Optionally, integrating the set of modality values to be tested to obtain a unique modality value, and updating the optimized modality key-value pair with the unique modality value to obtain the final modality key-value pair, includes:

[0039] Calculate the similarity between every two modality values to be tested in the set of modality values to be tested, and obtain one or more modality value similarities;

[0040] If one or more modality value similarities are all greater than a preset modality similarity, randomly extract one modality value to be tested from the multiple modality values to be tested, and confirm the extracted modality value to be tested as the unique modality value, and replace the value of the optimized modality key-value pair with the unique modality value to obtain the final modality key-value pair;

[0041] Otherwise, update the set of modality values to be tested to a preset empty set state to obtain an empty modality value to be tested, use the empty modality value to be tested as the set of modality values to be tested, and return to the step of receiving the target original data based on the optimized modality key-value pair.

[0042] Optionally, indexing and coding the official information using a pre-constructed coding method to obtain product identification data, includes:

[0043] Obtain multiple initial field values based on the official information, and use a pre-constructed hash function and a preset first length to calculate the first hash value of each initial field value in the multiple initial field values to obtain a first hash value set, where the initial field value corresponds to the unit modality information one by one, the first hash value corresponds to the initial field value one by one, and the length of the first hash value is the first length;

[0044] Use a pre-constructed index code calculation formula and the first hash value set to calculate the initial identification code corresponding to the luxury product, where the index code calculation formula is as follows:

[0045]

[0046]

[0047] Among them, S represents the initial identification code, SM 3 (*) represents a commercial password hash function, Z 1 represents the 1st initial field value among the multiple initial field values, Z 2 represents the 2nd initial field value among the multiple initial field values, Z i represents the ith initial field value among the multiple initial field values, Z n represents the nth initial field value among the multiple initial field values, n represents that there are n initial field values in total, l 1 represents the first length, KDF(*) represents a password derivation function, l 2Represents a second length, M represents a first hash value set parameter, y q Represents the q-th first hash value in the first hash value set;

[0048] Based on the luxury set, summarize the initial identification codes to obtain an initial identification code library, and obtain the product identification data corresponding to the official information based on the initial identification code library.

[0049] Optionally, the obtaining the product identification data corresponding to the official information based on the initial identification code library includes:

[0050] Extract the initial identification code corresponding to the luxury from the initial identification code library, where the initial identification codes in the initial identification code library correspond one by one to the luxuries in the luxury set;

[0051] If there is an initial identification code in the initial identification code library that is the same as the extracted initial identification code, randomly extract one initial field value from the multiple initial field values corresponding to the luxury, and copy the extracted initial field value, summarize the copied initial field value and the multiple initial field values to obtain an updated field value set, use the updated field value set as the multiple initial field values, and return to the step of calculating the first hash value of each initial field value in the multiple initial field values using the pre-constructed hash function and the preset first length until there is no initial identification code in the initial identification code library that is the same as the extracted initial identification code, and obtain the product identification data.

[0052] Optionally, constructing an authorization information structure tree and an anti-counterfeiting image structure tree based on the pre-constructed authorization data set and the pre-constructed anti-counterfeiting image set, constructing an authorization information structure tree and an anti-counterfeiting image structure tree based on the pre-constructed authorization data set and the pre-constructed anti-counterfeiting image set, includes:

[0053] Receive the anti-counterfeiting image set, where the anti-counterfeiting image set includes a genuine product detail image set and a forged detail image set;

[0054] Perform the following operations on each genuine product detail image in the genuine product detail image set:

[0055] Calculate the image similarity between the genuine product detail image and each forged detail image in the forged detail image set to obtain multiple image similarity values, extract the image similarity values greater than the preset image similarity threshold from the multiple image similarity values to obtain a high image similarity value set, if the high image similarity value set is not an empty set, obtain one or more forged detail images corresponding to the high image similarity value set to obtain a high imitation image set, and use the pre-constructed detail identifier to associate the high imitation image set and the genuine product detail image to obtain unit anti-counterfeiting image data, where the unit anti-counterfeiting image data includes a detail identifier, an associated high imitation image set, and a genuine product detail image;

[0056] Using the product identification data as the parent node and the detail identification in the unit anti-counterfeiting image data as the child node, construct the unit anti-counterfeiting image path, and summarize the unit anti-counterfeiting image paths to obtain the anti-counterfeiting image structure tree;

[0057] Extract the authorization data from the authorization dataset in sequence, where the authorization data includes the store identification and store information, and the store information includes: store name, store address, legal person, and authorization time;

[0058] Using the product identification data as the parent node and the store identification in the authorization data as the child node, construct the unit authorization information path, and summarize the unit authorization information paths to obtain the authorization information structure tree.

[0059] To achieve the above object, the present invention also provides a luxury goods authentication database construction system based on multi-modal features, including:

[0060] An official information extraction module, used to obtain a set of luxury goods, and perform the following operations on each luxury good in the set of luxury goods: extract official information based on the luxury good, where the official information includes multiple unit modal information, and the multiple unit information includes: brand, item number, place of origin, luxury good name, initial production year and month, release batch, limited quantity information, and single-piece style;

[0061] A product identification data acquisition module, used to perform index code encoding on the official information using a pre-constructed encoding method to obtain product identification data;

[0062] A structure tree construction module, used to sequentially extract unit modal information from the official information, identify the target modal key based on the extracted unit modal information, use the product identification data as the parent node and the target modal key as the child node to construct the product-unit modal path, summarize the product-unit modal paths to obtain the official information structure tree, and construct the authorization information structure tree and the anti-counterfeiting image structure tree based on the pre-constructed authorization dataset and the pre-constructed anti-counterfeiting image set;

[0063] An integration module, used to integrate the authorization information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree to obtain the unit modal database corresponding to the luxury good, and summarize the unit modal databases to obtain the luxury goods authentication database based on multi-modal features.

[0064] To solve the above problems, the present invention also provides an electronic device, and the electronic device includes:

[0065] A memory, storing at least one instruction;

[0066] A processor, executing the instructions stored in the memory to implement the above-mentioned luxury goods authentication database construction method based on multi-modal features.

[0067] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is executed by a processor in an electronic device to implement the above-mentioned method for constructing a luxury goods identification database based on multimodal features.

[0068] In order to solve the problem described in the background technology, the present invention first obtains a set of luxury goods, and performs the following operations on each luxury good in the set of luxury goods: extracting official information based on the luxury goods, wherein the official information includes multiple unit modal information, and the multiple unit information includes: brand, item number, place of origin, name of the luxury goods, initial production year and month, launch batch, limited edition information and single style. The embodiment of the present invention establishes an initial official information extraction model to eliminate irrelevant or invalid information, retains the valid information required in the initial official information extraction model, and adds standard reference objects to the construction of a luxury goods database by constructing key-value pairs and a structure tree. The official information is indexed and encoded using a pre-constructed encoding method to obtain product identification data. In the present invention, when an initial identification code identical to the initial identification code exists in the initial identification code library, the calculation result of the index code calculation formula is changed by changing the first hash value set, so that the initial identification code uniquely corresponds to the luxury goods and is confirmed as product identification data. Extract unit modal information from official information in sequence, identify target modal key based on extracted unit modal information, construct product-unit modal path with product identification data as parent node and target modal key as child node, summarize product-unit modal path, obtain official information structure tree, construct authorization information structure tree and anti-counterfeiting image structure tree based on pre-constructed authorization data set and pre-constructed anti-counterfeiting image set, integrate the authorization information structure tree, official information structure tree and anti-counterfeiting image structure tree, obtain unit modal database corresponding to luxury goods, the embodiment of the present invention integrates official information, authorization information and anti-counterfeiting image, and stores them in unit modal database in the form of structure tree, summarizes unit modal database, obtains luxury goods identification database based on multi-modal features. The present invention establishes corresponding database for each luxury goods in the luxury goods collection, and the corresponding database includes official information, sales channels, appearance features of known counterfeit luxury goods and detailed features of genuine goods of the luxury goods, so when it is necessary to identify and verify the anti-counterfeiting of the luxury goods, it can be compared by calling the data in the database, so as to perform anti-counterfeiting identification of the luxury goods. Therefore, the present invention can realize the construction of database of official information of luxury goods and counterfeit luxury goods according to multi-modal data features. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 A schematic diagram of a flow chart of a method for constructing a luxury goods identification database based on multimodal features provided in one embodiment of the present invention;

[0070] Figure 2 Function module diagram of a luxury goods authentication database construction system based on multi-modal features provided by an embodiment of the present invention;

[0071] Figure 3 Structural schematic diagram of an electronic device for implementing the method for constructing a luxury goods authentication database based on multi-modal features provided by an embodiment of the present invention;

[0072] Explanation of reference numerals:

[0073] 1. Electronic device; 10. Processor; 11. Memory; 12. Bus.

[0074] The implementation, functional features and advantages of the objectives of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0075] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0076] An embodiment of the present application provides a method for constructing a luxury goods authentication database based on multi-modal features. The execution subject of the method for constructing a luxury goods authentication database based on multi-modal features includes but is not limited to at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for constructing a luxury goods authentication database based on multi-modal features can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.

[0077] Refer to Figure 1 As shown, it is a flow schematic diagram of a method for constructing a luxury goods authentication database based on multi-modal features provided by an embodiment of the present invention. In this embodiment, the method for constructing a luxury goods authentication database based on multi-modal features includes:

[0078] S1. Obtain a luxury goods set, and perform the following operations on each luxury good in the luxury goods set: Extract official information based on the luxury good, where the official information includes multiple unit modal information, and the multiple unit information includes: brand, item number, place of origin, luxury good name, initial production year and month, release batch, limited edition information, and single-piece style.

[0079] It can be understood that the luxury goods set is a set of multiple luxury goods constructed artificially. Exemplarily, if a database administrator deems a certain product to be a luxury good, then the certain product is confirmed as a luxury good in the luxury goods set.

[0080] In the embodiments of the present invention, a corresponding database is established for each luxury item concentrated in luxury items. The corresponding database includes official information, sales channels, appearances of known counterfeit luxury items, and appearances of genuine products of the luxury item. Therefore, when it is necessary to identify and verify the authenticity of the luxury item, the database can be called to provide a reference for the anti-counterfeiting identification of the luxury item. Therefore, the embodiments of the present invention aim to provide a method for constructing a database integrating official information of luxury items and information related to counterfeit luxury item data that can be used for anti-counterfeiting identification.

[0081] It should be noted that the luxury item name is the name set by the manufacturer of the luxury item for the luxury item. For example, a certain perfume of brand A is called Bluebell, and at this time, Bluebell is the luxury item name. The initial production date is the date when the manufacturer of the luxury item first conducts production. For example, if product B is first sold on the date of a year, month, and day of a, b, and c, and if it is found that the production date of a product B held by a certain user is before the initial production date, then product B can be identified as a counterfeit luxury item.

[0082] It can be understood that the launch batch is the corresponding production batch of the luxury item. The limited quantity information includes whether the luxury item is a limited luxury item. If it is a limited luxury item, it also includes the limited quantity. The single-piece style includes the style information of the luxury item. For example, a certain product C has single-piece styles including blue, red, and green, and the single-piece styles corresponding to different luxury items in the luxury item concentration are different.

[0083] It should be noted that in the embodiments of the present invention, the luxury item is a set composed of multiple unit luxury items with exactly the same official information. The unit luxury item refers to a product. For example, if there are two basketballs with exactly the same official information, then both basketballs are the luxury items. If there is any different unit modal information in the official information corresponding to two luxury items, they can be considered as two different luxury items in the luxury item concentration. Therefore, each unit modal information in the official information is uniquely determined.

[0084] It should be noted that in actual use, the official information includes but is not limited to the multiple unit modal information. When the multiple unit modal information cannot meet the usage requirements, other unit modal information can be adaptively added to expand the database. For example: if the luxury item is a down jacket, the filling amount can also be added, and the method of adaptively adding other unit modal information is the same as the method of adding the multiple unit modal information.

[0085] It can be understood that the official information is the information provided by the manufacturer and brand party of the luxury item, used to distinguish it from the information of unauthorized manufacturers and brand parties.

[0086] Furthermore, the extraction of official information based on luxury items includes:

[0087] Acquire multiple single-modal storage definitions based on multiple unit modal information, wherein the unit modal information corresponds to the single-modal storage definition one by one, and the single-modal storage definition includes: field name, field annotation, field type and field constraint condition;

[0088] Extracting unimodal storage definitions from multiple unimodal storage definitions in sequence, extracting field names based on the extracted unimodal storage definitions, and constructing initial modal key-value pairs with the field names as keys and pre-constructed blank values ​​and field annotations as values, wherein the key in the initial modal key-value pairs is the field name, the value in the initial modal key-value pairs is the field annotation and the blank value, and the data type of the blank value is the field type;

[0089] The initial modal key-value pairs are summarized to obtain an initial modal key-value pair set, the initial modal key-value pair set is confirmed as an initial official information extraction model of the luxury product, and official information is extracted based on the initial official information extraction model.

[0090] It is understandable that the field annotation is an explanation of the field name, and the present invention does not limit the encoding method of the field name. For example, if the field name is encoded in English, such as brand, the field annotation is an annotation that can be understood by users or administrators when querying, such as brand. The field type is the data type when storing unit modal information, including but not limited to integer type, text type, time type, etc.

[0091] It should be noted that the key-value pair is a basic structure for data storage, consisting of a key and a value, the key acting as a unique identifier, and the value being the specific information associated with the key. Through the key, the corresponding value can be found quickly and accurately, and this technology is a prior art. In an embodiment of the present invention, the initial modal key-value pair is a key-value pair constructed with the field name as the key and the pre-constructed blank value and field annotation as the value. For the convenience of illustration, the embodiment of the present invention uses text to briefly describe the initial modal key-value pair. For example, the initial modal key value includes: [Key: field name; Value: field annotation, blank value]. The initial modal key-value pair set is a set consisting of multiple initial modal key-value pairs, and corresponds one-to-one to multiple unit modal information. Exemplarily, if the unit modal information corresponding to the initial modal key value is a brand, the initial modal key value is: [Key: brand; Value: brand, Null], and Null represents a blank value.

[0092] It should be noted that the initial official information extraction model is used to screen the information received from the official website or manufacturer, etc. Since there may be a large amount of irrelevant or invalid information when receiving information, the embodiment of the present invention establishes an initial official information extraction model to eliminate irrelevant or invalid information and retain the valid information required in the initial official information extraction model.

[0093] Furthermore, extracting official information based on the initial official information extraction model includes:

[0094] Receiving product original data, parsing the product original data to obtain product structured data and product unstructured data, where the product structured data includes multiple unit structure data;

[0095] Sequentially extracting unit structure data from the product structured data, identifying the structure primary key and structure value in the unit structure data, identifying the primary key character segment corresponding to the structure primary key, and using a pre-constructed semantic similarity algorithm to calculate the character similarity between the field name corresponding to each initial modal key-value pair in the initial modal key-value pair set and the primary key character segment, obtaining a character similarity set;

[0096] Extracting the character similarities greater than a preset similarity value from the character similarity set to obtain a high-similarity character set. If the high-similarity character set is an empty set, the unit structure data is excluded;

[0097] Otherwise, extracting the maximum character similarity from the high-similarity character set, converting the structure value using the field type to obtain a filled value corresponding to the field type, and mapping the filled value to the blank value of the initial modal key-value pair corresponding to the maximum character similarity to obtain a filled modal key-value pair;

[0098] Aggregating the filled modal key-value pairs to obtain multiple filled modal key-value pairs, and using the multiple filled modal key-value pairs to update the initial official information extraction model to obtain a sub-optimal official information extraction model;

[0099] Using the product unstructured data to update the sub-optimal official information extraction model to obtain an optimized official information extraction model, and extracting official information based on the optimized official information extraction model.

[0100] It can be understood that the product original data is information received from an official website or a manufacturer, etc., such as: tables, images, process cards, operation cards, etc. The purpose of parsing the product original data is to distinguish the product structured data and product unstructured data in the product original data. This technology is an existing technology and will not be elaborated here. The product structured data is the structured data screened from the product original data. In the embodiments of the present invention, the data structure of the structured data is a key-value pair, consisting of a key and a value. The key and value of the unit structure data are the structure primary key and structure value respectively. The primary key character segment is the field name represented by the structure primary key. For example, the product structured data includes: [Key: s2ikgh; Value: perfume], where the primary key character segment is the field name [s2ikgh] represented by [Key: s2ikgh].

[0101] It is understandable that the semantic similarity algorithm is an algorithm for calculating the semantic similarity between two texts. In the prior art, including recognition models based on deep learning, word vector models, etc., all can achieve the effects in the embodiments of the present invention, and the embodiments of the present invention are not limited herein.

[0102] Specifically, the character similarity set is a set composed of multiple character similarities. The character similarity is the similarity between the field name of the initial modal key-value pair calculated by using the semantic similarity algorithm and the primary key character segment. The preset similarity value is a value set manually for screening the character similarities in the character similarity set. When the character similarity is greater than the preset similarity value, it is considered that the field name of the initial modal key-value pair corresponding to the character similarity is similar to the primary key character segment. If the high similarity character set is an empty set, it means that the unit structure data is not the content that the initial official information extraction model wants to extract. Therefore, the unit structure data needs to be excluded. When the high similarity character set is not an empty set, it means that there is one or more character similarities in the high similarity character set that are greater than the preset similarity value. In the embodiments of the present invention, the initial modal key-value pair corresponding to the maximum character similarity is determined as the object to be filled in the unit structure data. Therefore, the structure value in the unit structure data is filled into the blank value of the initial modal key-value pair corresponding to the maximum character similarity. However, before filling, the structure value needs to be transformed so that the structure value conforms to the single-modal storage definition of the initial modal key-value pair corresponding to the maximum character similarity.

[0103] It is understandable that updating the initial official information extraction model by using multiple filling modal key-value pairs means: using multiple filling modal key-value pairs to replace the corresponding initial modal key-value pairs in the initial official information extraction model.

[0104] Exemplarily, the initial modal key-value pair corresponding to the maximum character similarity in the high similarity character set is [Key: brand; Value: brand, Null], where Null represents the blank value, and the unit structure data is [Structure primary key: brand; Structure value: Louis]. Then, the structure value in the unit structure data needs to be transformed into a filling value that meets the field type of the corresponding initial modal key-value pair. If the filling value is [Louis], then the filling value [Louis] is mapped to the blank value of [Key: brand; Value: brand, Null] to obtain the filling modal key-value pair [Key: brand; Value: brand, Louis].

[0105] Further, the initial official information extraction model includes: the initial modal key-value pair [key: brand; value: brand, Null], and the corresponding filling modal key-value pair is [key: brand; value: brand, Louis]. Updating the initial official information extraction model with the filling modal key-value pair means replacing [key: brand; value: brand, Null] in the initial official information extraction model with [key: brand; value: brand, Louis]. That is, only the filling modal key-value pair [key: brand; value: brand, Louis] exists in the sub-optimal official information extraction model.

[0106] Further, updating the sub-optimal official information extraction model with the product unstructured data to obtain an optimized official information extraction model, and extracting official information based on the optimized official information extraction model includes:

[0107] Extracting initial text data and image data including text from the product unstructured data, identifying the text in the image data including text to obtain image text, and summarizing the image text and the initial text data to obtain the text to be analyzed;

[0108] Performing keyword classification based on entity extraction, relationship extraction, and attribute extraction on the text to be analyzed based on the sub-optimal official information extraction model to obtain a keyword group set, where the keyword group set includes multiple keyword groups, and each keyword group includes a main word text and one or more content word texts;

[0109] Converting the keyword groups in the keyword group set into key-value pairs to obtain multiple keyword group key-value pairs, summarizing the multiple keyword group key-value pairs to obtain transformed structured data, and updating the sub-optimal official information extraction model with the transformed structured data to obtain an optimized official information extraction model, where the optimized official information extraction model includes multiple optimized modal key-value pairs, and the keyword group key-value pair is a key-value pair formed with the main word text as the key and the one or more content word texts as the value;

[0110] Performing data cleaning and integrity verification on the optimized official information extraction model, and obtaining a complete official information model after confirming that the optimized official information extraction model after performing data cleaning and integrity verification is a preset complete model, and extracting official information based on the complete official information model.

[0111] It should be noted that the product unstructured data is the data in the product original data except for the product structured data. The initial text data is the text data in the product unstructured data, the image data including text is the image data in the product unstructured data, and the image data includes text. The technology of identifying the text in the image data including text to obtain image text is an existing technology, and the data format of the image text is the same as that of the initial text.

[0112] It is understandable that the operation of performing keyword classification based on entity extraction, relationship extraction, and attribute extraction on the text to be analyzed is a prior art and a basic task in natural language processing (NLP). The aim is to identify keywords with specific meanings from the text to be analyzed and classify the keywords into predefined categories. In the embodiments of the present invention, the operation of performing keyword classification based on entity extraction, relationship extraction, and attribute extraction on the text to be analyzed is carried out to obtain a keyword group set.

[0113] Furthermore, the conversion of the keyword group set into key-value pairs to obtain the converted structured data is as follows: The main word text of each keyword group in the keyword group set is used as the key, and one or more content word texts in the keyword group are used as the values and then summarized to obtain the converted structured data. For example, the keyword group key-value pair is: [Key: main word text; Value: content word text, content word text].

[0114] It is understandable that the main word text in the keyword group corresponds to the text that can be used as the key in the converted structured data, and one or more content word texts correspond to the text that can be used as the value in the converted structured data. That is, when performing the operation of keyword classification based on entity extraction, relationship extraction, and attribute extraction, the corresponding predefined category is the character segment in the initial modal key-value pair in the initial official information extraction model.

[0115] It should be noted that the operation of updating the sub-optimal official information extraction model with the converted structured data to obtain the optimized official information extraction model is similar to the method of updating the initial official information extraction model with the product structured data. However, the difference is that when updating the initial official information extraction model with the product structured data, each initial modal key-value pair in the initial official information extraction model has a blank value, while in the sub-optimal official information extraction model, due to the existence of filled modal key-value pairs and there is no blank value in the filled modal key-value pairs. Therefore, when updating the filled modal key-value pairs with the data in the converted structured data, it is to add values to the filled modal key-value pairs.

[0116] Exemplarily, if there are respectively a filling mode key-value pair [key: brand; value: brand, Louis] and an initial mode key-value pair [key: weight; value: weight, blank value] in the sub-optimal official information extraction model, and among them, the converted structured data is: [key: brand; value: Louis], then calculate the character similarity between the converted structured data in the example and the initial mode key-value pair and the filling mode key-value pair respectively. Finally, it is determined that the converted structured data can be used to update the filling mode key-value pair [key: brand; value: brand, Louis]. Therefore, the initial mode key-value pair is updated with the structured data, so that the optimized mode key-value pair is: [key: brand; value: brand, Louis, Louis].

[0117] Further, after performing data cleaning and integrity verification on the optimized official information extraction model, and confirming that the optimized official information extraction model after performing data cleaning and integrity verification is the preset complete model, a complete official information model is obtained, including:

[0118] Sequentially extract the optimized mode key-value pairs from the multiple optimized mode key-value pairs corresponding to the optimized official information extraction model, and perform the following operations on the extracted optimized mode key-value pairs:

[0119] Identify the values of the optimized mode key-value pairs to obtain an optimized mode value set. Among them, the optimized mode value set includes one or more optimized mode values. When there is one optimized mode value in the optimized mode value set, the optimized mode value is a field annotation. When there are multiple optimized mode values in the optimized mode value set, the multiple optimized mode values include the field annotation, and also include the values in the filling mode key-value pairs corresponding to the optimized mode key-value pairs, or the values in the keyword group key-value pairs corresponding to the optimized mode key-value pairs;

[0120] Remove the optimized mode values corresponding to the field annotations from the optimized mode value set, and summarize the remaining optimized mode values to obtain a to-be-verified mode value set. Determine whether the to-be-verified mode value set meets the field constraint conditions corresponding to the optimized mode key-value pairs. Among them, the optimized mode key-value pairs correspond one-to-one with the initial mode key-value pairs, the field constraint conditions corresponding to the optimized mode key-value pairs are the same as the field constraint conditions of the corresponding initial mode key-value pairs, and the character constraint conditions of all initial mode key-value pairs are: unique value and not null value;

[0121] If the to-be-verified mode value sets corresponding to each optimized mode key-value in the multiple optimized mode keys all meet the field constraint conditions, then confirm that the optimized official information extraction model is the preset complete model, and obtain the complete official information model;

[0122] If the set of to-be-verified modal values corresponding to the extracted optimized modal key-value pairs does not meet the field constraint conditions and the set of to-be-verified modal values is an empty set, then receive the target original data based on the optimized modal key-value pairs, parse the target original data, use the parsed target original data to fill the set of to-be-verified modal values to obtain a filled modal value set, and use the filled modal value set to update the optimized modal key-value pairs. Take the updated optimized modal key-value pairs as the optimized modal key-value pairs, and return the step of sequentially extracting optimized modal key-value pairs from the multiple optimized modal key-value pairs corresponding to the optimized official information extraction model;

[0123] If the set of to-be-verified modal values does not meet the field constraint conditions and the set of to-be-verified modal values is not an empty set, then confirm that the set of to-be-verified modal values includes multiple to-be-verified modal values, integrate the set of to-be-verified modal values to obtain a unique modal value, and use the unique modal value to update the optimized modal key-value pairs to obtain the final modal key-value pairs;

[0124] Summarize the final modal key-value pairs to obtain a complete official information model.

[0125] Exemplarily, the extracted optimized modal key-value pairs are: [Key: brand; Value: brand, Louis, Louis], then the optimized modal value set is [brand, Louis, Louis]. Remove the optimized modal value [brand] corresponding to the field annotation from the optimized modal value set, and summarize the remaining optimized modal values [Louis] and [Louis] to obtain the set of to-be-verified modal values [Louis, Louis]. In this example, there are multiple optimized modal values in the optimized modal value set [brand, Louis, Louis], which are: brand, Louis, and Louis. Among them, "brand" is a field annotation, and "Louis" and "Louis" can be filled by the values in the filled modal key-value pairs or the values in the keyword group key-value pairs.

[0126] It should be noted that in the invention embodiment, the initial official information extraction model includes: an initial set of modal key-value pairs, and the values in the initial modal key-value pairs include: field annotations and blank values. That is, in the initial modal key-value pairs, storage space is reserved for blank values.

[0127] Furthermore, use the product structured data to update the initial official information extraction model to obtain a sub-optimal official information extraction model. Therefore, the sub-optimal official information extraction model includes: multiple initial modal key-value pairs (the initial modal keys that have not been updated by the filled modal key-value pairs are still retained) and multiple filled modal key-value pairs (the initial modal key-value pairs updated by the filled modal key-value pairs), and the total number of multiple initial modal key-value pairs and multiple filled modal key-value pairs is the same as the number of initial modal key-value pairs in the initial official information extraction model.

[0128] Further, by using multiple keyword group key-value pairs, transformed structured data is obtained, and after updating the sub-optimal official information extraction model with the transformed structured data, regardless of whether the multiple initial modal key-value pairs and multiple filling modal keys in the original sub-optimal official information extraction model are updated, they are all called optimized modal key-value pairs in the optimized official information extraction model. Therefore, the total number of optimized modal key-value pairs in the optimized official information extraction model is the same as the number of initial modal key-value pairs in the initial official information extraction model.

[0129] Further, when constructing key-value pairs such as initial modal key-value pairs, optimized modal key-value pairs, and filling modal key-value pairs in the embodiments of the present invention, one piece of data is a unit value in the value of the key-value pair. For example, when updating the sub-optimal official information extraction model with multiple keyword group key-value pairs, if the keyword group key-value pair is: [Key: main word text; Value: content word text, content word text], that is, the multiple content word texts are split into two unit values: [content word text, content word text]. Further, in the optimized modal value set, the [content word text, content word text] corresponds to two optimized modal values.

[0130] It can be understood that the optimized modal value set is a set composed of all values in the corresponding optimized modal key-value pairs. The unique value in the character constraint condition, and not being a null value means that there is exactly one to-be-tested modal value in the to-be-tested modal value set, and the to-be-tested modal value is not a null value. For example, the to-be-tested modal value set [Louis, Louis] does not meet the "unique value", and the to-be-tested modal value set [Null] does not meet the "not being a null value".

[0131] Further, if the to-be-tested modal value set is an empty set, it means that in the previous text of the embodiments of the present invention, the values in the optimized modal key-value pairs have not been effectively filled, and there is a problem of missing data. Therefore, it is necessary to fill the values in the optimized modal key-value pairs. Thus, the embodiments of the present invention receive target original data, and the target original data is data received based on the missing values in the optimized modal key-value pairs. For example, if the optimized modal key-value pair is: [Key: weight; Value: weight, blank value], then the target original data is data received based on weight. For example, asking the manufacturer about the weight of luxury goods.

[0132] It can be understood that the method of parsing the target original data and using the parsed target original data to fill the to-be-tested modal value set to obtain the filled modal value set is the same as the method of parsing the product original data and obtaining the optimized modal value set in the optimized modal key-value pairs through the parsed product original data, that is, the same as the method of processing the product original data in the embodiments of the present invention, and can achieve the same effect. The filled modal value set is the value in the optimized modal key-value pairs obtained by using the target original data.

[0133] Exemplarily, by parsing the target original data, a filled modal value set of [1 kg, 1 kg] can be obtained. The original optimized modal key-value pair is: [Key: weight; Value: weight, blank value]. Then, the optimized modal key-value pair is updated using the filled modal value set: the blank value in the original optimized modal key-value pair [Key: weight; Value: weight, blank value] is replaced with [1 kg, 1 kg], resulting in the updated optimized modal key-value pair [Key: weight; Value: weight, 1 kg, 1 kg].

[0134] It should be noted that the data cleaning in the embodiments of the present invention is to integrate the to-be-verified modal value set.

[0135] Furthermore, integrating the to-be-verified modal value set to obtain a unique modal value, and updating the optimized modal key-value pair using the unique modal value to obtain the final modal key-value pair includes:

[0136] Calculating the similarity between every two to-be-verified modal values in the to-be-verified modal value set to obtain one or more modal value similarities;

[0137] If one or more modal value similarities are all greater than a preset modal similarity, randomly extract one to-be-verified modal value from the multiple to-be-verified modal values, and confirm the extracted to-be-verified modal value as the unique modal value. Use the unique modal value to replace the value of the optimized modal key-value pair to obtain the final modal key-value pair;

[0138] Otherwise, update the to-be-verified modal value set to a preset empty set state to obtain an empty to-be-verified modal value. Use the empty to-be-verified modal value as the to-be-verified modal value set, and return to the step of receiving the target original data based on the optimized modal key-value pair.

[0139] It can be understood that the method of calculating the similarity between multiple to-be-verified modal values in the to-be-verified modal value set is the same as the method of calculating the character similarity between the field name corresponding to each initial modal key-value pair in the initial modal key-value pair set and the primary key character segment using a pre-constructed semantic similarity algorithm, and can achieve the same effect, which will not be elaborated here.

[0140] Specifically, the preset modal similarity is similar to the definition of the preset similarity value, and is used to judge the similarity degree between multiple to-be-verified modal values. If one or more modal value similarities are all greater than the preset modal similarity, it indicates that the to-be-verified modal values in the to-be-verified modal value set are basically similar and have the same semantics, and can all be used. Therefore, in the embodiments of the present invention, one to-be-verified modal value is randomly extracted, and the extracted to-be-verified modal value is confirmed as the unique modal value. The unique modal value is used as the value in the final modal key-value pair, that is, the data finally stored in the official information structure tree.

[0141] Further, if it does not meet the condition that the similarity of the one or more modal values is greater than the preset modal similarity, it indicates that the differences in semantics and text among multiple to-be-tested modal values in the to-be-tested modal value set may be caused by error information or text damage. Therefore, in the embodiment of the present invention, the to-be-tested modal value set is confirmed to be in an empty set state, that is, all the content in the to-be-tested modal value set is deleted to obtain an empty to-be-tested modal value, and the empty to-be-tested modal value is used as the to-be-tested modal value set, and the step of receiving the target original data based on the optimized modal key-value pair is returned to re-extract information from the to-be-tested value set.

[0142] S2. Index code encode the official information by using a pre-constructed encoding method to obtain product identification data.

[0143] Further, the index code encoding the official information by using a pre-constructed encoding method to obtain product identification data includes:

[0144] Obtain multiple initial field values based on the official information, and use a pre-constructed hash function and a preset first length to calculate the first hash value of each initial field value in the multiple initial field values to obtain a first hash value set, where the initial field value corresponds to the unit modal information one by one, the first hash value corresponds to the initial field value one by one, and the length of the first hash value is the first length;

[0145] Use a pre-constructed index code calculation formula and the first hash value set to calculate the initial identification code corresponding to the luxury product, where the index code calculation formula is as follows:

[0146]

[0147]

[0148] Among them, S represents the initial identification code, SM 3 (*) represents a commercial password hash function, Z 1 represents the first initial field value among the multiple initial field values, Z 2 represents the second initial field value among the multiple initial field values, Z i represents the i-th initial field value among the multiple initial field values, Z n represents the n-th initial field value among the multiple initial field values, n represents that there are n initial field values in total, l 1 represents the first length, KDF(*) represents a password derivation function, l 2 represents the second length, M represents the first hash value set parameter, y q represents the q-th first hash value in the first hash value set;

[0149] The initial identification codes are aggregated based on the luxury goods collection to obtain an initial identification code library, and the product identification data corresponding to the official information is obtained based on the initial identification code library.

[0150] Furthermore, the initial field value is the value of the unit modal information in the official information. For example, there is a final modal key-value pair: [key: weight; value: weight, 1kg]. According to the previous definition of this embodiment, the value is [1kg], and the corresponding initial field value is [1kg].

[0151] It is understandable that the first hash value is a hash value of a first length obtained by calculating the initial field value using the commercial cryptographic hash function. The first length is used to limit the length of the hash value generated by the commercial cryptographic hash function to avoid excessively long hash values ​​that increase the burden on the computer.

[0152] Furthermore, the password derivation function and the commercial password hash function are prior art and are not described in detail here. The second length is similar to the first length and is used to limit the length of the hash value calculated by the password derivation function. The first hash value set parameter is a parameter for adding the first hash value set to the initial identification code.

[0153] It is understandable that, in the embodiment of the present invention, a corresponding initial identification code can be calculated for each luxury item in the luxury item collection, and the initial identification code library can be obtained by aggregating these initial identification codes.

[0154] Furthermore, the obtaining of product identification data corresponding to the official information based on the initial identification code library includes:

[0155] Extracting an initial identification code corresponding to the luxury goods from an initial identification code library, wherein the initial identification codes in the initial identification code library correspond one-to-one to the luxury goods in the luxury goods collection;

[0156] If there exists an initial identification code identical to the extracted initial identification code in the initial identification code library, then an initial field value is randomly extracted from the multiple initial field values ​​corresponding to the luxury product, and the extracted initial field value is copied, the copied initial field value and the multiple initial field values ​​are aggregated to obtain an updated field value set, so that the updated field value set is the multiple initial field values, and the step of calculating the first hash value of each of the multiple initial field values ​​using a pre-built hash function and a preset first length is returned, until there is no initial identification code identical to the extracted initial identification code in the initial identification code library, and the product identification data is obtained.

[0157] It is understandable that if there is an initial identification code in the initial identification code library that is the same as the initial identification code, that is, at least two different luxury goods have the same initial identification code. At this time, the extracted initial identification code cannot uniquely identify the corresponding luxury good. Therefore, in the embodiments of the present invention, an initial field value is extracted, and the extracted initial field value is copied. The copying and extraction methods are prior arts and will not be elaborated herein.

[0158] It is understandable that when there is an initial identification code in the initial identification code library that is the same as the initial identification code, by copying a random initial field value set in the initial field value set, the multiple initial field values are updated, so as to achieve the purpose of changing the number of first hash values in the first hash value set, thereby realizing the change of the calculation result of the index code calculation formula until there is no initial identification code in the initial identification code library that is the same as the initial identification code, that is, until the initial identification code corresponds uniquely to the luxury good, then the initial identification code is confirmed as the product identification data.

[0159] S3. Sequentially extract unit modal information from the official information, and identify the target modal key based on the extracted unit modal information.

[0160] Furthermore, in the embodiments of the present invention, when obtaining the official information based on the complete official information model, the final modal key-value pair in the complete official model is used as the official information. Therefore, when sequentially extracting unit modal information from the official information and identifying the target modal key based on the extracted unit modal information, the target modal key is the key in the corresponding final modal key-value pair.

[0161] S4. Using the product identification data as the parent node and the target modal key as the child node, construct a product-unit modal path, and summarize the product-unit modal paths to obtain the official information structure tree.

[0162] It should be noted that using the product identification data as the parent node and the target modal key as the child node to construct a product-unit modal path is a method for constructing a path between two data, and this method is a conventional method for constructing a path in a structure tree and is a prior art, which will not be elaborated herein.

[0163] It is understandable that the product identification data is the unique identifier of the luxury good, and the target modal key is the key in the final modal key-value pair. Therefore, the constructed product-unit modal path is [product identification data - target modal key], and the target modal key is connected to the value in the final modal key-value pair. Therefore, in the official information structure tree, the value in the corresponding final modal key-value pair can be queried through [product identification data - target modal key].

[0164] Further, in the official information structure tree obtained by summarizing the product-unit modal path, since the product identification data is the parent node, when summarizing the product-unit modal path, the corresponding product identification data is the parent node, and each different target modal key is a child node of the product identification data. In the embodiment of the present invention, this data structure is called the official information structure tree.

[0165] S5. Construct an authorized information structure tree and an anti-counterfeiting image structure tree based on a pre-constructed authorized data set and a pre-constructed anti-counterfeiting image set.

[0166] Further, the constructing of the authorized information structure tree and the anti-counterfeiting image structure tree based on the pre-constructed authorized data set and the pre-constructed anti-counterfeiting image set includes:

[0167] Receive an anti-counterfeiting image set, where the anti-counterfeiting image set includes a genuine product detail image set and a forged detail image set;

[0168] Perform the following operations on each genuine product detail image in the genuine product detail image set:

[0169] Calculate the image similarity between the genuine product detail image and each forged detail image in the forged detail image set to obtain a plurality of image similarity values. Extract the image similarity values greater than a preset image similarity threshold from the plurality of image similarity values to obtain a high image similarity value set. If the high image similarity value set is not an empty set, obtain one or more forged detail images corresponding to the high image similarity value set to obtain a high imitation image set. Use the pre-constructed detail identifier to associate the high imitation image set and the genuine product detail image to obtain unit anti-counterfeiting image data, where the unit anti-counterfeiting image data includes a detail identifier, an associated high imitation image set, and a genuine product detail image;

[0170] Construct a unit anti-counterfeiting image path with the product identification data as the parent node and the detail identifier in the unit anti-counterfeiting image data as the child node, and summarize the unit anti-counterfeiting image paths to obtain an anti-counterfeiting image structure tree;

[0171] Sequentially extract authorized data from the authorized data set, where the authorized data includes a store identifier and store information, and the store information includes: store name, store address, legal person, and authorization time;

[0172] Construct a unit authorized information path with the product identification data as the parent node and the store identifier in the authorized data as the child node, and summarize the unit authorized information paths to obtain an authorized information structure tree.

[0173] It is understandable that the genuine product detail atlas consists of images of various details of genuine luxury goods produced by the official, and the images have sufficient light sources and can be used as image references for identifying counterfeit luxury goods. For example: genuine brand logo images, genuine zipper images, etc. The counterfeit detail atlas consists of various detail images of the collected counterfeit luxury goods, including: counterfeit brand logo images, counterfeit zipper images, etc.

[0174] Specifically, the method of calculating the image similarity between the genuine product detail images and each counterfeit detail image in the counterfeit detail atlas to obtain multiple image similarity values is a prior art. For example, calculating the similarity of images based on pixel-level differences, using deep learning feature extraction to calculate image similarity, and there are multiple prior arts that can achieve this. The embodiments of the present invention do not limit this here.

[0175] It is understandable that the image similarity threshold is a value set artificially for screening images with a high degree of similarity to the genuine product detail images from the counterfeit detail atlas. Therefore, the high-quality imitation atlas includes one or more counterfeit detail images, all of which are images similar to the genuine product detail images. For example, if the genuine product detail image is a genuine brand logo image, the high-quality imitation atlas may be one or more counterfeit brand logo images. And use the pre-constructed detail identifier [brand logo image] to associate the high-quality imitation atlas and the genuine product detail images to obtain unit anti-counterfeiting image data. That is, the detail identifier is a character used to identify the high-quality imitation atlas and the genuine product detail images, and the unit anti-counterfeiting image data is the high-quality imitation atlas and the genuine product detail images after detail identification. The association is to store both the high-quality imitation atlas and the genuine product detail images after detail identification in the unit anti-counterfeiting image data.

[0176] Furthermore, the method of constructing the unit anti-counterfeiting image path with the product identification data as the parent node and the detail identifier in the unit anti-counterfeiting image data as the child node is the same as the method of constructing the product-unit modality path with the product identification data as the parent node and the target modality key as the child node, and can achieve the same effect, which will not be elaborated here. Summarizing the unit anti-counterfeiting image paths to obtain the anti-counterfeiting image structure tree is the same as the method of summarizing the product-unit modality paths to obtain the official information structure tree, and can achieve the same effect, which will not be elaborated here.

[0177] It is understandable that the authorized dataset is data obtained from the official authorized stores. The store identifier is the unique identifier constructed by the authorized store. Therefore, the method of constructing the unit authorized information path with the product identification data as the parent node and the store identifier in the authorized data as the child node, and summarizing the unit authorized information paths to obtain the anti-counterfeiting image structure tree is the same as the method of constructing the unit anti-counterfeiting image path with the product identification data as the parent node and the detail identifier in the unit anti-counterfeiting image data as the child node, and summarizing the unit anti-counterfeiting image paths to obtain the anti-counterfeiting image structure tree, and can achieve the same effect, which will not be elaborated here.

[0178] S6. Integrate the authorized information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree to obtain the unit modal database corresponding to the luxury good.

[0179] It can be understood that in the three different structure trees of the authorized information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree, the product identification data is used as the parent node for construction. When integrating the authorized information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree to obtain the unit modal database corresponding to the luxury good, if the three different structure trees of the authorized information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree are directly unified under the unique product identification data, it will increase the time consumption when respectively hoping to retrieve one aspect of the authorized information, the official information, and the anti-counterfeiting image. Therefore, the embodiments of the invention respectively use the pre-constructed official information product structure data, the pre-constructed authorized information product structure data, and the pre-constructed anti-counterfeiting image product structure data to replace the corresponding product identification data in the authorized information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree, and use the product identification data as the parent node, and respectively use the official information product structure data, the authorized information product structure data, and the anti-counterfeiting image product structure data as the child nodes to construct the structure tree for integration, so as to obtain the unit modal database corresponding to the luxury good.

[0180] Exemplarily, before integration, the official information structure tree is [Product Identification Data - Official Information], the authorized information structure tree is [Product Identification Data - Authorized Information], and the anti-counterfeiting image structure tree is [Product Identification Data - Anti-counterfeiting Image]. After integration, the official information structure tree is [Official Information Product Identification Data - Official Information], the authorized information structure tree is [Authorized Information Product Identification Data - Authorized Information], and the anti-counterfeiting image structure tree is [Anti-counterfeiting Image Product Identification Data - Anti-counterfeiting Image], and the official information structure tree, the authorized information structure tree, and the anti-counterfeiting image structure tree are integrated through the product identification data, and the official information product identification data, the authorized information product identification data, and the anti-counterfeiting image product identification data are all child nodes of the product identification data.

[0181] S7. Summarize the unit modal database to obtain the luxury good identification database based on multi-modal features.

[0182] It can be understood that the luxury good identification database based on multi-modal features is the sum of multiple unit modal databases.

[0183] To solve the problems described in the background art, the present invention first obtains a luxury set, and performs the following operations on each luxury in the luxury set: extracting official information based on the luxury, where the official information includes multiple unit modal information, and the multiple unit information includes: brand, item number, place of origin, luxury name, initial production date, release batch, limited quantity information, and single-piece style. The embodiment of the present invention establishes an initial official information extraction model to eliminate irrelevant or invalid information, retains the valid information required in the initial official information extraction model, and adds a standard reference object to the construction of the luxury database in the form of constructing key-value pairs and structure trees. Using a pre-constructed coding method to encode the official information into index code, product identification data is obtained. In the present invention, when there is an initial identification code in the initial identification code library that is the same as the initial identification code, the calculation result of the index code calculation formula is changed by changing the first hash value set, so that the initial identification code corresponds uniquely to the luxury, and it is confirmed as product identification data. Sequentially extract unit modal information from the official information, identify the target modal key based on the extracted unit modal information, use the product identification data as the parent node, and the target modal key as the child node to construct a product-unit modal path, summarize the product-unit modal path, obtain the official information structure tree, construct an authorization information structure tree and an anti-counterfeiting image structure tree based on the pre-constructed authorization data set and the pre-constructed anti-counterfeiting image set, integrate the authorization information structure tree, the official information structure tree and the anti-counterfeiting image structure tree, obtain the unit modal database corresponding to the luxury. The embodiment of the present invention integrates official information, authorization information and anti-counterfeiting images, and stores them in the unit modal database in the form of a structure tree, summarizes the unit modal database, and obtains a luxury identification database based on multi-modal features. The present invention establishes a corresponding database for each luxury in the luxury set, and the corresponding database includes the official information, sales channels, known appearance features of counterfeit luxuries, and detailed features of genuine products of the luxury. Therefore, when it is necessary to identify and verify the anti-counterfeiting of the luxury, the data in the database can be called for comparison, so as to conduct anti-counterfeiting identification of the luxury. Therefore, the present invention can realize the construction of a database of luxury official information and counterfeit luxuries according to multi-modal data features.

[0184] As Figure 2 shown, it is a functional module diagram of a luxury identification database construction system based on multi-modal features provided by an embodiment of the present invention.

[0185] The luxury goods authentication database construction system 100 based on multi-modal features according to the present invention can be installed in an electronic device. According to the functions achieved, the luxury goods authentication database construction system 100 based on multi-modal features can include an official information extraction module 101, a product identification data acquisition module 102, a structure tree construction module 103, and an integration module 104. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, and are stored in the memory of the electronic device.

[0186] The official information extraction module 101 is used to obtain a set of luxury goods and perform the following operations on each luxury good in the set of luxury goods: extract official information based on the luxury good, where the official information includes multiple unit modal information, and the multiple unit information includes: brand, item number, place of origin, luxury good name, initial production date, release batch, limited quantity information, and single-piece style;

[0187] The product identification data acquisition module 102 is used to perform index code encoding on the official information using a pre-constructed encoding method to obtain product identification data;

[0188] The structure tree construction module 103 is used to sequentially extract unit modal information from the official information, identify target modal keys based on the extracted unit modal information, use the product identification data as the parent node, and the target modal keys as the sub-nodes to construct a product-unit modal path, summarize the product-unit modal paths to obtain an official information structure tree, and construct an authorized information structure tree and an anti-counterfeiting image structure tree based on a pre-constructed authorized data set and a pre-constructed anti-counterfeiting image set;

[0189] The integration module 104 is used to integrate the authorized information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree to obtain a unit modal database corresponding to the luxury good, and summarize the unit modal databases to obtain a luxury goods authentication database based on multi-modal features.

[0190] Specifically, each module in the luxury goods authentication database construction system 100 based on multi-modal features in the embodiment of the present invention uses the same technical means as the Figure 1 luxury goods authentication database construction method based on multi-modal features described above and can produce the same technical effects, which will not be elaborated here.

[0191] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing a luxury goods authentication database construction method based on multi-modal features provided by an embodiment of the present invention.

[0192] The electronic device 1 may include a processor 10, a memory 11, and a bus 12. It may also include a computer program stored in the memory 11 and executable on the processor 10, such as a program for building a luxury goods authentication database based on multimodal features.

[0193] Among them, the memory 11 includes at least one type of readable storage medium, which includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 also includes the internal storage unit of the electronic device 1 and also includes external storage devices. The memory 11 can not only be used to store application software installed on the electronic device 1 and various types of data, such as the code of the program for building a luxury goods authentication database based on multimodal features, but can also be used to temporarily store data that has been output or will be output.

[0194] In some embodiments, the processor 10 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits. By running or executing programs or modules stored in the memory 11 (such as the program for building a luxury goods authentication database based on multimodal features, etc.), and by calling data stored in the memory 11, it performs various functions of the electronic device 1 and processes data.

[0195] The bus 12 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is configured to implement the connection and communication between the memory 11 and at least one processor 10, etc.

[0196] Figure 3 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 3 The shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have a different component layout.

[0197] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management system, so as to implement functions such as charge management, discharge management, and power consumption management through the power management system. The power source may also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0198] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0199] Optionally, the electronic device 1 may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0200] The program of the method for constructing a luxury goods authentication database based on multi-modal features stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve the following:

[0201] Obtain a set of luxury goods, and perform the following operations on each luxury good in the set of luxury goods:

[0202] Extract official information based on the luxury goods. Among them, the official information includes multiple unit modal information, and the multiple unit information includes: brand, item number, origin, luxury good name, initial production year and month, release batch, limited quantity information, and single-piece style;

[0203] Use a pre-constructed coding method to perform index code coding on the official information to obtain product identification data;

[0204] Extract unit modal information from the official information in sequence, identify the target modal key based on the extracted unit modal information, use the product identification data as the parent node, and the target modal key as the child node to construct a product-unit modal path, summarize the product-unit modal path, and obtain an official information structure tree;

[0205] Construct an authorized information structure tree and an anti-counterfeiting image structure tree based on a pre-constructed authorized data set and a pre-constructed anti-counterfeiting image set;

[0206] Integrate the authorized information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree to obtain a unit modal database corresponding to the luxury goods;

[0207] Summarize the unit modal database to obtain a luxury goods authentication database based on multi-modal features.

[0208] Specifically, for the specific implementation method of the above instructions by the processor 10, reference can be made to Figures 1 to 3 The description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0209] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or system that can carry the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0210] The present invention also provides a computer-readable storage medium. The readable storage medium stores a computer program. When the computer program is executed by the processor of the electronic device, it can achieve:

[0211] Obtain a luxury set, and perform the following operations on each luxury in the luxury set:

[0212] Extract official information based on the luxury. The official information includes multiple unit modal information, and the multiple unit information includes: brand, item number, place of origin, luxury name, initial production date, release batch, limited edition information, and single-piece style;

[0213] Use a pre-constructed coding method to encode the official information into index codes to obtain product identification data;

[0214] Extract unit modal information from the official information in sequence, identify the target modal key based on the extracted unit modal information, use the product identification data as the parent node, and the target modal key as the child node to construct a product-unit modal path, and summarize the product-unit modal paths to obtain an official information structure tree;

[0215] Construct an authorized information structure tree and an anti-counterfeiting image structure tree based on a pre-constructed authorized data set and a pre-constructed anti-counterfeiting image set;

[0216] Integrate the authorized information structure tree, the official information structure tree, and the anti-counterfeiting image structure tree to obtain a unit modal database corresponding to the luxury;

[0217] Summarize the unit modal database to obtain a luxury identification database based on multi-modal features.

[0218] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative, and there may be other partitioning methods in actual implementation.

[0219] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0220] In addition, the functional modules in each embodiment of the present invention can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0221] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0222] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for constructing a luxury goods identification database based on multimodal features, characterized in that: The method comprises: Get the luxury goods set and perform the following operations on each luxury item in the luxury goods set: Extract official information based on luxury goods, where the official information includes multiple unit modal information, and the multiple unit information includes: brand, item number, origin, luxury goods name, initial production year and month, launch batch, limited edition information and single piece style; Use a pre-built encoding method to index the official information to obtain product identification data; Extract unit modal information from the official information in sequence, identify the target modal key based on the extracted unit modal information, construct a product-unit modal path with the product identification data as the parent node and the target modal key as the child node, summarize the product-unit modal path, and obtain the official information structure tree; Constructing an authorization information structure tree and an anti-counterfeiting image structure tree based on a pre-constructed authorization data set and a pre-constructed anti-counterfeiting image set; Integrate the authorization information structure tree, the official information structure tree and the anti-counterfeiting image structure tree to obtain a unit modality database corresponding to luxury goods; The unit modal database is aggregated to obtain a luxury goods identification database based on multimodal features.

2. The method for constructing a luxury goods identification database based on multimodal features according to claim 1, characterized in that: The official information extracted based on luxury goods includes: Acquire multiple single-modal storage definitions based on multiple unit modal information, wherein the unit modal information corresponds to the single-modal storage definition one by one, and the single-modal storage definition includes: field name, field annotation, field type and field constraint condition; Extracting unimodal storage definitions from multiple unimodal storage definitions in sequence, extracting field names based on the extracted unimodal storage definitions, and constructing initial modal key-value pairs with the field names as keys and pre-constructed blank values ​​and field annotations as values, wherein the key in the initial modal key-value pairs is the field name, the value in the initial modal key-value pairs is the field annotation and the blank value, and the data type of the blank value is the field type; The initial modal key-value pairs are summarized to obtain an initial modal key-value pair set, the initial modal key-value pair set is confirmed as an initial official information extraction model of the luxury product, and official information is extracted based on the initial official information extraction model.

3. The method for constructing a luxury goods identification database based on multimodal features according to claim 2, characterized in that: The extracting official information based on the initial official information extraction model includes: Receiving product original data, parsing the product original data, and obtaining product structured data and product unstructured data, wherein the product structured data includes a plurality of unit structure data; Extracting unit structure data from product structured data in sequence, identifying the structure primary key and structure value in the unit structure data, identifying the primary key character segment corresponding to the structure primary key, and using a pre-built semantic similarity algorithm to calculate the character similarity between the field name corresponding to each initial modal key-value pair in the initial modal key-value pair set and the primary key character segment to obtain a character similarity set; Extracting character similarities greater than a preset similarity value from the character similarity set to obtain a high-similar character set, and if the high-similar character set is an empty set, removing the unit structure data; Otherwise, extract the maximum character similarity from the high-similarity character set, convert the structure value using the field type to obtain a fill value corresponding to the field type, map the fill value to the blank value of the initial modal key-value pair corresponding to the maximum character similarity, and obtain a fill modal key-value pair; Summarizing the filled modal key-value pairs to obtain a plurality of filled modal key-value pairs, and using the plurality of filled modal key-value pairs to update the initial official information extraction model to obtain a suboptimal official information extraction model; The suboptimal official information extraction model is updated using product unstructured data to obtain an optimized official information extraction model, and official information is extracted based on the optimized official information extraction model.

4. The method for constructing a luxury goods identification database based on multimodal features according to claim 3, characterized in that: The method of updating the suboptimal official information extraction model by using the product unstructured data to obtain an optimized official information extraction model, and extracting official information based on the optimized official information extraction model, includes: Extracting initial text data and image data including text from product unstructured data, identifying text in the image data including text to obtain image text, and summarizing the image text and initial text data to obtain text to be analyzed; Based on the suboptimal official information extraction model, keyword classification based on entity extraction, relationship extraction and attribute extraction is performed on the text to be analyzed to obtain a keyword group set, wherein the keyword group set includes multiple keyword groups, and each keyword group includes a main word text and one or more content word texts; Convert the keyword group in the keyword group set into a key-value pair to obtain a plurality of keyword group key-value pairs, summarize the plurality of keyword group key-value pairs to obtain converted structured data, and use the converted structured data to update the suboptimal official information extraction model to obtain an optimized official information extraction model, wherein the optimized official information extraction model includes a plurality of optimized modal key-value pairs, and the keyword group key-value pair is a key-value pair consisting of a main word text as a key and the one or more content word texts as a value; Perform data cleaning and integrity check on the optimized official information extraction model, confirm that the optimized official information extraction model after data cleaning and integrity check is the preset complete model, obtain the complete official information model, and extract official information based on the complete official information model.

5. The method for constructing a luxury goods identification database based on multimodal features according to claim 4, characterized in that: The step of performing data cleaning and integrity check on the optimized official information extraction model and confirming that the optimized official information extraction model after data cleaning and integrity check is a preset complete model, thereby obtaining a complete official information model, includes: Extract optimization modal key-value pairs in sequence from multiple optimization modal key-value pairs corresponding to the optimization official information extraction model, and perform the following operations on the extracted optimization modal key-value pairs: Identify the value of the optimized modal key-value pair to obtain an optimized modal value set, wherein the optimized modal value set includes one or more optimized modal values, when there is one optimized modal value in the optimized modal value set, the optimized modal value is a field annotation, when there are multiple optimized modal values ​​in the optimized modal value set, the multiple optimized modal values ​​include the field annotation and also include a value in a padding modal key-value pair corresponding to the optimized modal key-value pair, or a value in a keyword group key-value pair corresponding to the optimized modal key-value pair; Eliminate the optimized modal values ​​corresponding to the field annotations from the optimized modal value set, and summarize the retained optimized modal values ​​to obtain the modal value set to be tested, and determine whether the modal value set to be tested meets the field constraints corresponding to the optimized modal key-value pairs, where the optimized modal key-value pairs correspond one-to-one to the initial modal key-value pairs, the field constraints corresponding to the optimized modal key-value pairs are the same as the field constraints of the corresponding initial modal key-value pairs, and the character constraints of all initial modal key-value pairs are: unique values ​​and not empty values; If the set of modal values ​​to be checked corresponding to each optimized modal key value in the multiple optimized modal key values ​​satisfies the field constraint condition, then the optimized official information extraction model is confirmed to be a preset complete model, and a complete official information model is obtained; If the modal value set to be tested corresponding to the extracted optimized modal key-value pair does not meet the field constraint condition, and the modal value set to be tested is an empty set, then based on the optimized modal key-value pair, the target original data is received, the target original data is parsed, and the modal value set to be tested is filled with the parsed target original data to obtain the filled modal value set, and the optimized modal key-value pair is updated with the filled modal value set, and the updated optimized modal key-value pair is used as the optimized modal key-value pair, and the step of sequentially extracting optimized modal key-value pairs from the multiple optimized modal key-value pairs corresponding to the optimized official information extraction model is returned; If the modal value set to be checked does not meet the field constraint condition, and the modal value set to be checked is not an empty set, confirming that the modal value set to be checked includes multiple modal values ​​to be checked, integrating the modal value set to be checked to obtain a unique modal value, and using the unique modal value to update the optimized modal key-value pair to obtain a final modal key-value pair; Summarize the final modal key-value pairs to obtain the complete official information model.

6. The method for constructing a luxury goods identification database based on multimodal features according to claim 5, characterized in that: The step of integrating the set of modal values ​​to be tested to obtain a unique modal value, and updating the optimized modal key-value pair using the unique modal value to obtain a final modal key-value pair includes: Calculate the similarity between every two modal values ​​to be tested in the modal value set to be tested, and obtain one or more modal value similarities; If the similarities of one or more modal values ​​are greater than the preset modal similarity, a modal value to be tested is randomly extracted from the multiple modal values ​​to be tested, and the extracted modal value to be tested is confirmed as a unique modal value, and the value of the optimized modal key-value pair is replaced with the unique modal value to obtain a final modal key-value pair; Otherwise, the modal value set to be checked is updated to a preset empty set state to obtain an empty modal value to be checked, and the empty modal value to be checked is used as the modal value set to be checked, and the step of receiving the target original data based on the optimized modal key-value pair is returned.

7. The method for constructing a luxury goods identification database based on multimodal features according to claim 6, characterized in that: The method of using the pre-built encoding method to index the official information to obtain product identification data includes: Acquire multiple initial field values ​​based on official information, and calculate a first hash value of each of the multiple initial field values ​​using a pre-constructed hash function and a preset first length to obtain a first hash value set, wherein the initial field value corresponds to the unit modal information one-to-one, the first hash value corresponds to the initial field value one-to-one, and the length of the first hash value is the first length; The initial identification code corresponding to the luxury item is calculated using the pre-built index code calculation formula and the first hash value set, wherein the index code calculation formula is as follows: Wherein, S represents an initial identification code, SM3(*) represents a commercial cryptographic hash function, Z1 represents the first initial field value among multiple initial field values, Z2 represents the second initial field value among multiple initial field values, and Z i Represents the i-th initial field value among multiple initial field values, Z n represents the nth initial field value among multiple initial field values, n represents a total of n initial field values, l1 represents the first length, KDF(*) represents the password derivation function, l2 represents the second length, M represents the first hash value set parameter, y q represents the qth first hash value in the first hash value set; The initial identification codes are aggregated based on the luxury goods collection to obtain an initial identification code library, and the product identification data corresponding to the official information is obtained based on the initial identification code library.

8. The method for constructing a luxury goods identification database based on multimodal features according to claim 7, characterized in that: The obtaining the product identification data corresponding to the official information based on the initial identification code library includes: Extracting an initial identification code corresponding to the luxury goods from an initial identification code library, wherein the initial identification codes in the initial identification code library correspond one-to-one to the luxury goods in the luxury goods collection; If there exists an initial identification code identical to the extracted initial identification code in the initial identification code library, then an initial field value is randomly extracted from the multiple initial field values ​​corresponding to the luxury product, and the extracted initial field value is copied, the copied initial field value and the multiple initial field values ​​are aggregated to obtain an updated field value set, so that the field value set is updated to the multiple initial field values, and the step of calculating the first hash value of each of the multiple initial field values ​​using a pre-built hash function and a preset first length is returned, until there is no initial identification code identical to the extracted initial identification code in the initial identification code library, and the product identification data is obtained.

9. The method for constructing a luxury goods identification database based on multimodal features according to claim 8, characterized in that: The method of constructing an authorization information structure tree and an anti-counterfeiting image structure tree based on a pre-constructed authorization data set and a pre-constructed anti-counterfeiting image set comprises: receiving an anti-counterfeiting image set, wherein the anti-counterfeiting image set includes an authentic detail image set and a forged detail image set; For each authentic product detail image in the authentic product detail image set, perform the following operations: Calculate the image similarity between the genuine detail image and each forged detail image in the forged detail image set to obtain multiple image similarity values, extract the image similarity value greater than a preset image similarity threshold from the multiple image similarity values ​​to obtain a high image similarity value set, if the high image similarity value set is not an empty set, obtain one or more forged detail images corresponding to the high image similarity value set to obtain a high-imitation image set, use a pre-constructed detail identifier to associate the high-imitation image set with the genuine detail image to obtain unit anti-counterfeiting image data, wherein the unit anti-counterfeiting image data includes the detail identifier, the associated high-imitation image set and the genuine detail image; Taking the product identification data as the parent node and the detail identification in the unit anti-counterfeiting image data as the child node to construct a unit anti-counterfeiting image path, summarizing the unit anti-counterfeiting image paths to obtain an anti-counterfeiting image structure tree; Extracting authorization data from the authorization data set in sequence, wherein the authorization data includes a store identifier and store information, wherein the store information includes: store name, store address, legal person and authorization time; The product identification data is used as a parent node, and the store identification in the authorization data is used as a child node to construct a unit authorization information path, and the unit authorization information path is summarized to obtain an authorization information structure tree.

10. A luxury goods identification database construction system based on multimodal features, characterized in that: The system comprises: The official information extraction module is used to obtain a set of luxury goods, and performs the following operations on each luxury good in the set of luxury goods: extracting official information based on the luxury goods, wherein the official information includes multiple unit modal information, and the multiple unit information includes: brand, item number, origin, name of the luxury goods, initial production year and month, launch batch, limited edition information and single piece style; A product identification data acquisition module, used to encode the official information with an index code using a pre-built encoding method to obtain product identification data; A structure tree construction module is used to extract unit modal information from official information in sequence, identify the target modal key based on the extracted unit modal information, construct a product-unit modal path with product identification data as the parent node and the target modal key as the child node, summarize the product-unit modal path, obtain the official information structure tree, and construct an authorization information structure tree and an anti-counterfeiting image structure tree based on a pre-constructed authorization data set and a pre-constructed anti-counterfeiting image set; The integration module is used to integrate the authorization information structure tree, the official information structure tree and the anti-counterfeiting image structure tree to obtain the unit modality database corresponding to the luxury goods, summarize the unit modality database, and obtain the luxury goods identification database based on multimodal features.