Data object semantic identifier generation method and device
By introducing tree structure and multi-level classification information optimization codebook construction in RQ-VAE, the problem of poor semantic recognition in the classification organization of data objects in the prior art is solved, and more accurate data object classification and retrieval is achieved.
Patent Information
- Application Number
- CN202510474490.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
Smart Images

Figure CN120336589A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data identification, and particularly relates to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating semantic identifiers of data objects. Background Art
[0002] A data identification system is a system for uniquely identifying digital objects, usually with characteristics such as uniqueness, standardization, and persistence. In the digital age, a sound data identification system is the basis for realizing the efficient organization and cross-domain circulation of large-scale data resources. Its importance is reflected in: First, it avoids data confusion and duplication by assigning globally unique data identifiers to avoid conflicts between data from different sources; second, it can form a set of data organization systems to efficiently manage large-scale data resources; third, it improves the data transfer efficiency, can optimize data mapping and rapid identification and matching, and enhance the automation level of data exchange. Generally speaking, data identification plays a key role in data circulation and use, and can be used to classify, categorize, and organize data, facilitating the identification, search, and retrieval of data to improve the efficiency of data management.
[0003] There are many existing data identification coding methods, and the mainstream identifications are represented by DOI (Digital Object Identifier) identification and CSTR (China Science and Technology Resource) identification.
[0004] 1) DOI Digital Object Identifier. DOI, also known as Digital Object Identifier, is an internationally recognized identification system. It is a standard for digital resources, providing a unique and persistent identifier or handle for various objects, with the characteristics of permanently naming and dynamically resolving links for resources. It is developed by the International Organization for Standardization (ISO). DOI is also applicable to the Uniform Resource Identifier (URI) system. They are widely used to identify academic and professional information, such as journal articles, research reports, datasets, and official publications. After the DOI identification number is resolved, it can be linked to one or more pieces of information. However, the identification number itself has no relation to the information obtained after resolution, and it may not be possible to obtain all the information, only the information of relevant publications. The DOI resolution protocol can be found in RFC3652, RFC 3651 describes the naming mechanism, and RFC 3650 describes its architecture. In actual applications, DOI is mostly resolved through websites. An actual example of DOI is a paper on a certain website, whose DOI is 10.48550 / arXiv.2412.06381. Among them, 10.48550 is the unique number assigned by DOI to this website. The website can assign unique IDs to each paper by itself, thus forming a unique ID for this paper. When browsing the website https: / / doi.org / 10.48550 / arXiv.2412.06381, the paper information corresponding to the identification number can be seen. This depends on the support of a set of resolution service systems in the background.
[0005] 2) CSTR Identifier. The CSTR identifier system is developed by the Computer Network Information Center of the Chinese Academy of Sciences and is an important carrier for China's national standards to serve global scientific and technological resources. The Scientific and Technological Resources Identification Service Platform is built based on the national standard GB / T 32843—2016 "Identification of Scientific and Technological Resources" and provides a unique identification service for global scientific data, papers, preprints, patents and other scientific and technological resources. For example, the SciEngine platform of Science Press has been connected to CSTR, enabling resources such as scientific papers to be identified and cited through CSTR.
[0006] These two identification systems are currently mainly applicable to the open sharing of scientific and technological resources and rapid positioning and acquisition on a global scale. They have uniqueness and persistence. Semantic identification means that on the basis of the existing uniqueness, the identification realizes semantic representation. There have been some preliminary explorations on semantic identification in the field of generative retrieval in the past two years. Academic researchers have proposed several semantic identification generation methods, mainly divided into title-based identification and semantic structured identification methods.
[0007] 1) Title-based identification directly uses a data object such as the title of a document as its ID. The title is usually a short and rich summary of the entire document, providing a macro overview of the information contained within the document. Additionally, in some knowledge bases (such as Wikipedia), the title is usually unique, so it is an ideal choice as an identifier. This method has also been proven effective in knowledge-intensive language tasks. However, the drawback is that not all documents have high-quality titles.
[0008] 2) Semantically structured identification compresses the semantic expression of a data object such as a document into a shorter combination of numbers as the document semantic identifier. Its goal is to capture the semantic information of the document and automatically generate an identifier ID that can convey the semantic information of its corresponding document. The structure of the identifier DI is mainly used to effectively reduce the search space after each decoding step in generative retrieval. A typical method is to construct the ID of each document through kmeans clustering, which may share prefixes among semantically similar documents. As Figure 1 shown, another typical method is to generate the identifier ID through the compression quantization method of RQ-VAE (Residual Quantized Variational Autoencoder).
[0009] Specifically, as Figure 1 shown, for the original data (such as articles, images, videos), we assume that each data object has relevant content features that can capture useful semantic information (such as titles, descriptions, or images). Additionally, assume that we can access a pre-trained content encoder to generate a semantic embedding x ∈ R d . For example, general pre-trained text encoders (such as Sentence-T5 and BERT) can be used to transform the text features of the project to obtain semantic embeddings. Then, the semantic embeddings are quantized by RQ-VAE to generate semantic IDs for each data object.
[0010] The basic principle of using RQ-VAE to implement semantic identification is as follows: Define the semantic ID as a tuple of codewords of length m. Each codeword in the tuple comes from a different codebook. Therefore, the number of items that can be uniquely represented by the semantic ID is equal to the product of the codebook sizes. Although different techniques for generating semantic IDs may result in IDs with different semantic attributes, we hope they at least have the following property: Similar items (items with similar content features or semantically close embeddings) should have overlapping semantic IDs. For example, an item with semantic ID (10, 21, 35) should be more similar to an item with semantic ID (10, 21, 40) than to an item with ID (10, 23, 32). Next, we discuss the quantization scheme for semantic ID generation. RQ-VAE is a multi-level vector quantizer that quantizes the residuals to generate a tuple of codewords (also known as semantic IDs). The autoencoder is jointly trained by updating the quantization codebook and the DNN encoder parameters.
[0011] First, the input x is encoded by the encoder E to learn the latent representation z = ε(x). At layer 0 (d = 0), the initial residual is simply defined as r0 = z. At each layer d, we have a codebook where k is the size of the codebook. Then, r0 is quantized by mapping it to the nearest embedding in the codebook of that layer. The index of the nearest embedding at d = 0, i.e., c0 = argmin i ||r0 - e k ||, represents the 0th codeword. For the next layer d = 1, the residual is defined as Then, similar to level 0, the code of the first level is calculated by finding the embedding in the codebook that is closest to r1. This process is recursively repeated m times to obtain a tuple consisting of m codewords, and these codewords represent the formation of the semantic identification. This recursive method approximates the input from coarse to fine granularity. Note that we choose to use a separate codebook of size K for each of the m levels instead of using a single codebook of size K. This is because the standard of the residuals decreases as the level increases, thus allowing different granularities at different levels.
[0012] Identification methods such as DOI and CSTR only have uniqueness and persistence, but do not have semantics. They can only be used to specifically locate a certain data object by using this identifier, and the premise is that this identifier is required. While semantic identification can judge the similarity of data objects through the similarity of semantic identifiers between two data objects, or like the semantic representation of RQ-VAE, it can also organize semantic information from coarse granularity to fine granularity.
[0013] The problem with the identification method of RQ-VAE is that the representation space of data objects is the same for each type of data, and each type of data object reuses the semantics of all codes starting from the second codebook.
[0014] In real life, for different data objects, such as for commodities, the size differences of different commodity classification systems are huge. Some commodities can be classified into multiple levels and multiple types, while some commodities can only be classified into small categories (with a small number of categories). Therefore, the existing RQ-VAE is not applicable to real scenarios.
[0015] In summary, in the prior art, not considering semantic information or considering semantic information but adopting the unified RQVAE codebook reuse method will lead to the problems that the data distribution space sizes are the same and the codebook reuse results in poor semantic recognition. This defect is caused by not considering the large differences in the classification and organization systems of each data object in the real scenario, and the semantic identification space should be differentiated. Summary of the Invention
[0016] The purpose of the present invention is to solve the problem of poor semantic recognition of the above-mentioned prior art, and at the same time not considering the problem of unbalanced data distribution in the real scenario, and proposes a new method and system for generating semantic identification of data objects.
[0017] Aiming at the deficiencies of the prior art, as Figure 5 shown, the present invention proposes a method for generating semantic identification of data objects, which includes:
[0018] Initial step, obtaining the target data x to be identified and constructing a tree-shaped codebook model. The structure of this model is a tree, and each node in the tree represents a codebook; it is set that the level corresponding to the root node of this tree is the 0th level;
[0019] Encoding step, extracting the latent representation z of the target data through an encoder;
[0020] Quantization step, at the 0th layer of the model, that is, the root node of the tree, the initial residual is r0 = z; map r0 to the nearest embedding in the node codebook to quantize r0, and the quantization result is an information representation value of the semantic identification; is used as the current residual; the embedding is used as the current embedding; this information representation value is used as the representation value of the current 0th level;
[0021] Intermediate step, processing the 1st layer of the tree-shaped codebook model, selecting the child node of the code corresponding to the representation value obtained in the previous step in this layer as the current codebook, and mapping the current residual r1 to the nearest embedding in the current codebook to quantize r1. And update the current residual r2 by subtracting the result of the current embedding from the updated current residual Another information representation value with the quantization result as the semantic identifier;
[0022] A loop step to update the current representation value with the another information representation value, and update the current embedding Update the current embedding, repeat the execution of this intermediate step n times again until reaching the codebook where the leaf node in the tree-shaped codebook is located, collect all the information representation values, and obtain a tuple composed of all the information representation values as the semantic identifier of the target data.
[0023] The method for generating the semantic identifier of a data object, which includes:
[0024] A decoding step, adding the quantization vectors of the codebooks to which the information representation values belong to obtain a summation result z', and the decoder decodes the summation result z' to obtain x', and constructs a loss function loss: where where sg represents the gradient stop operation, and r i is the updated current residual at the i-th execution of this intermediate step; train the encoder, the decoder, and the tree-shaped codebook through the loss function loss.
[0025] For the method for generating the semantic identifier of a data object in the case of a classification task, the loss function loss includes a classification loss Loss function loss:
[0026]
[0027] represents the cross-entropy loss of the k-th layer of the tree-shaped codebook; according to the category to which the target data belongs, obtain its semantic category id as the label, and according to the information representation value and the label of the k-th layer of the tree-shaped codebook, obtain the cross-entropy loss of the k-th layer.
[0028] The method for generating the semantic identifier of a data object, which includes:
[0029] A database building step, based on the target data with the semantic identifier, build a database, where the semantic identifier is used as the index of the target data, and the semantic identifier has the semantic category id to which the target data belongs;
[0030] A retrieval step, obtain the data to be retrieved, according to the semantics of the data to be retrieved, obtain the semantic id of the data to be retrieved, and retrieve the target data in the database that has the semantic id of the data to be retrieved as the retrieval result corresponding to the data to be retrieved.
[0031] As Figure 6 shown, the present invention also proposes a device for generating a semantic identifier of a data object, which includes:
[0032] An initial module that obtains target data x to be identified and constructs a tree-shaped codebook model, where each node of the tree-shaped codebook model represents a codebook; it is set that the level corresponding to the root node of the tree is the 0th level;
[0033] An encoding module that extracts the latent representation z of the target data through an encoder;
[0034] A quantization module, at the 0th layer of the tree-shaped codebook model, that is, the root node of the tree, the initial residual is r0 = z; maps r0 to the nearest embedding in the node codebook to quantize r0, and the quantization result is an information representation value of the semantic identifier; as the current residual; the embedding as the current embedding; and the information representation value as the representation value of the current 0th level;
[0035] An intermediate module that processes the 1st layer of the tree-shaped codebook model, selects the child node of the code corresponding to the representation value obtained in the previous step in this layer as the current codebook, and maps the current residual r1 to the nearest embedding in the current codebook to quantize r1. And update the current residual r2 by subtracting the result of the current embedding from the updated current residual The quantization result is another information representation value of the semantic identifier;
[0036] A loop module that updates the current representation value with the other information representation value, updates the current embedding with the embedding and repeats the execution of the intermediate module n times again until reaching the codebook where the leaf node in the tree-shaped codebook is located, and aggregates all the information representation values to obtain a tuple composed of all the information representation values as the semantic identifier of the target data.
[0037] The data object semantic identifier generation device, which includes:
[0038] A decoding module that adds up the quantization vectors of the codebooks to which the information representation values belong to obtain a summation result z ’ , and the decoder decodes the summation result z ’ to obtain x ’ , and constructs a loss function loss: where where sg represents the gradient stop operation, and r i is the current residual updated when the intermediate module is executed for the i-th time; trains the encoder, the decoder, and the tree-shaped codebook through the loss function loss.
[0039] For the data object semantic identifier generation device, for classification tasks, the loss function loss includes a classification loss Loss function loss:
[0040]
[0041] It represents the cross - entropy loss of the k - th layer of the tree - shaped codebook; according to the category to which the target data belongs, its semantic category id is obtained as the label, and based on the information representation value and the label of the k - th layer of the tree - shaped codebook, the cross - entropy loss of the k - th layer is obtained.
[0042] The database building module constructs a database based on the target data with the semantic identifier, where the semantic identifier is used as the index of the target data, and the semantic identifier has the semantic category id to which the target data belongs.
[0043] The retrieval module obtains the data to be retrieved, gets the semantic id of the data to be retrieved according to its semantics, and retrieves the target data in the database that has the semantic id of the data to be retrieved as the retrieval result corresponding to the data to be retrieved. Semantic ids with the same prefix indicate similar objects.
[0044] The present invention also proposes an electronic device, which includes the data object semantic identifier generation device described above. The electronic device is either connected to an information display device, and the information display device is used to display the semantic identifier of the target data with the display parameters, attributes set by the user or through an artificial intelligence model.
[0045] The present invention also proposes a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data object semantic identifier generation method are implemented.
[0046] The present invention also proposes a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the data object semantic identifier generation method are implemented.
[0047] As can be seen from the above solutions, the advantages of the present invention are as follows:
[0048] Compared with the existing technologies, in terms of semantic identifier generation, by making each semantic space independent, that is, having a separate codebook and a non - shared codebook strategy, the semantic identifier is made more distinguishable. By using the multi - level classification information of data objects as the supervision signal to optimize the construction of the codebook, the semantic identifier can more accurately capture the differences between different types of data, providing a novel method for organizing and managing data objects based on identifiers; through an iterative training scheme, the learning of each codebook can be accurately controlled, so as to better capture the semantic representation of data from coarse to fine granularity. Brief Description of the Drawings
[0049] Figure 1 Schematic diagram of RQVAE;
[0050] Figure 2 Semantic label generation method integrating tree structure and RQ-VAE method;
[0051] Figure 3 Flow chart of semantic label generation;
[0052] Figure 4 Tree structure diagram of semantic label encoding;
[0053] Figure 5 Flow chart of the method of the present invention;
[0054] Figure 6 Module diagram of the device of the present invention;
[0055] Figure 7 Schematic diagram of the structure of the first electronic device of the present invention;
[0056] Figure 8 Schematic diagram of the application environment structure of the first electronic device of the present invention;
[0057] Figure 9 Schematic diagram of the structure of the second electronic device of the present invention.
[0058] Reference signs:
[0059] A - First electronic device;
[0060] B - Data object semantic label generation device;
[0061] C - Data acquisition device;
[0062] D - Information display device;
[0063] 1000 - Second electronic device;
[0064] Ⅰ - Computing unit;
[0065] Ⅱ - ROM;
[0066] Ⅲ - RAM;
[0067] Ⅳ - Bus;
[0068] Ⅴ - Interface;
[0069] Ⅵ - Input unit;
[0070] Ⅶ - Output unit;
[0071] Ⅷ - Storage medium;
[0072] Ⅸ - Communication unit. Detailed implementation manners
[0073] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0074] Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0075] The processor described in the present invention is the control center of the electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0076] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0077] In a specific implementation, as an embodiment, the processor can include one or more CPUs. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device can include: servers, desktop computers, laptop computers, smartphones, tablets, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.
[0078] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0079] It should be noted that the structure of the electronic device shown in the accompanying drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than those shown in the drawings, or combine some components, or have different component arrangements.
[0080] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0081] It should also be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0082] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0083] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0084] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0085] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0086] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0087] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0088] When conducting research on semantic identification of data objects, the inventor found that in the prior art, if semantic information was not considered or if semantic information was considered but a unified RQVAE codebook reuse method was adopted, it would lead to problems such as the same size of the data distribution space and poor semantic recognition due to codebook reuse. This defect was caused by the failure to consider the significant differences in the classification and organization systems of various data objects in the real-world scenario, resulting in the need for differentiation in the semantic identification space. Through research on the existing identification method framework and classification and organization system, the inventor found that this defect could be solved by the RQVAE method based on a tree structure, that is, by adjusting the organization form of the original RQVAE linear codebook into a tree structure, so that each data object has its own independent semantic representation space when constructing semantic identification, and at the same time, the size and level of the codebook can be flexibly adjusted according to the different data volume distributions. The semantic identification obtained by the present invention is used to represent data. Especially from the perspective of classification, the prefixes of the semantic identification ids of data objects in the same category are the same or similar, which is conducive to subsequent classification, organization, retrieval, recommendation, etc. of the data object. When conducting retrieval, the corresponding semantic identification id can be quickly retrieved through the semantics of the text.
[0089] To achieve the above technical effects, the present invention proposes the following key technical points:
[0090] Key point 1, the main architecture of tree- and RQ-VAE-based semantic identification encoding; thus, it can flexibly adapt to various scenarios with significant differences in the distribution of multiple data objects, and at the same time, the semantic space of each codebook is independent, enabling the generated identification to be more recognizable;
[0091] Key point 2, using the multi-level classification information of data objects as a supervision signal to optimize the construction of the codebook; thus, it can strengthen the learning according to classification dimensions such as the field to which the data belongs, enabling the semantic identification to more accurately capture the differences between different types of data and providing a novel method for organizing and managing data objects based on identification;
[0092] Key point 3, an iterative training scheme. For the scenario where the data object is a multi-level and multi-label classification, we can more precisely control the codes of the codebooks at each layer in the tree structure through an iterative strategy, that is, each training can independently control the learning of one or the codebooks at the same layer in the tree structure; thus, it can more accurately capture the semantic differences from coarse to fine granularity of the data and the semantic differences at the same level. According to this difference, special model training can be carried out for the sub-classifications of each category respectively, so as to obtain a more accurate secondary classification. Improve the classification accuracy of semantic identification.
[0093] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and detailed descriptions are made in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.
[0094] The semantic identification coding framework proposed by the present invention is as follows Figure 2 As shown. The embedding on the left side is the same as the original RQ-VAE method, which is the semantic embedding obtained from the original data features through a general pre-trained text encoder (such as Sentence-T5 and BERT). Inside is the main structure of the semantic identification coding, which is a tree structure. Each node in the tree structure is a codebook, and the size of each codebook (i.e., the number of codes) can be flexibly set. The number of nodes in each layer of this tree structure is determined by the size of the codebook in the previous layer, and this size is controlled by hyperparameters. Through the main structure of this coding, we can finally obtain the quantized latent space semantic representation z, just like RQVAE. ’ , and then we use a DNN decoder to map the latent space semantic representation z ’ back to the semantic space x of the original data embedding ’ , with the aim of making the decoded semantic expression x ’ as consistent as possible with the original semantic expression x.
[0095] Regarding the intermediate core tree-shaped residual semantic coding structure. Our specific solution is as follows: Assume that the root node is the 0th layer, the children nodes of the root node are the 1st layer, and so on for subsequent nodes. represents the kth codebook (k starts from 0) in the lth layer. Then there are the following rule restrictions:
[0096] ● There is only 1 codebook in the 0th layer, i.e., the root node. The number of codes in this codebook is m, that is
[0097] m is controlled by hyperparameters.
[0098] ● The number of codebooks in the 1st to nth layers of the tree structure is the sum of the number of codes in all codebooks in the previous layer. Each codebook is respectively associated with one code in the upper layer. The number of codes in each codebook can be set separately.
[0099] ■ For example, for the first layer, the total number of its codebooks is m, and the number of codes in each codebook can be respectively set as
[0100] ■ For the second layer, the total number of its codebooks is ones.
[0101] ● Since each codebook (except for the codebook of the root node) is associated with only one code in the upper layer and is associated with the same number of codebooks as the number of codes in its own codebook in the lower layer, the structure is a multi-way tree structure.
[0102] ● When calculating the residual to obtain the semantic identifier, the algorithm logic is as follows:
[0103] S1. First, the encoder Encoder is also used to encode the embedding of the input x to learn the latent representation z = ε(x).
[0104] S2. At the 0th layer (d = 0), the initial residual is simply defined as r0 = z. r0 is quantized by mapping it to the nearest embedding in the codebook of this layer. The distance can be represented by calculating the similarity between r0 and each codeword in the codebook. The index of the nearest embedding at d = 0, that is, c0 = argmin i ||r0 - e k ||, representing the 0th codeword. If the codebook of the root node has m codes and r0 is closest to the 2nd code, then one information representation of the semantic identifier is 2. d represents the layer of the semantic identifier and is a preset value. For example, if there are a total of 4 layers of semantic granularity, then d = 4, which also means that the model will be trained from 4 codebooks.
[0105] S3. Enter the 1st layer, and the residual is defined as Since c0 = 2, the codebook of the child node associated with the 2nd code is located. Then, similar to the 0th level, the first-level code (i.e., the 2nd information representation of the semantic identifier) is calculated by finding the closest first-level embedding to r1 in the codebook.
[0106] S4. This process is recursively repeated n times until the codebook where the leaf node is located is traversed, thereby obtaining a tuple composed of n codewords, and these codewords represent the formed semantic identifier. This recursive method approximates the input from coarse to fine granularity.
[0107] S5. Finally, x is reconstructed. x is the representation after the original sample is embedded, |x| is a representation obtained through a calculation based on the codewords and then through a decompressor (which can be temporarily understood as a vector). The quantization vectors of the n codebooks where the codewords are located are added to obtain z ’ , and then x is obtained through the decoder ’ , where the quantization vector is the value of each code in the codebook. Our goal is to make the difference between x ’ and the original x as small as possible. The corresponding loss is: where Among them, sg represents the gradient stop operation, where β is a weight parameter. The meaning of sg is stop-gradient, that is, there is no gradient backpropagation, and g means there is gradient backpropagation. Both i and d refer to the codebook of the i-th layer. Through this loss, the encoder, decoder, and the intermediate tree-shaped codebook set can be jointly trained. During the training process, the number of layers and the size of the codebook remain unchanged, only the values of each code in the codebook change, which is equivalent to the parameters of the model.
[0108] In addition, starting from the field to which the data itself belongs is a good classification guidance. Therefore, the hierarchical classification labels of the data can be used as supervision signals to optimize the training process of semantic identification. Therefore, after the decoder, we perform supervised learning for classification through softmax. Since the semantic identification generated by the tree structure is a representation from coarse to fine granularity, we can perform classification supervision of multiple softmaxes based on multi-level classification labels, and control the weights through the parameter γ. Note: If γ is too large, the semantic identification encoding task can be directly mapped to a classification task, and the loss function of the classification task can be represented by cross-entropy loss, etc.
[0109]
[0110] The finally obtained loss function is:
[0111]
[0112] For the scenario where the data object is a multi-level and multi-label classification, this solution can also more finely control the codes of each layer of the codebook in the tree structure through an iterative strategy:
[0113] ● First, control the codebook of the 0th layer, and the loss function is as follows.
[0114]
[0115] ● After training, lock the codebook of the 0th layer and start training the codebook of the 1st layer. The loss function is
[0116]
[0117] ● After training, continue to lock the codebook of the 1st layer and start training the codebook of the 2nd layer. The loss function is as follows. And so on, each time lock the codebook of one layer and train the next layer until all codebooks are trained.
[0118]
[0119] Next, we will specifically use a specific example to illustrate the factual solution.
[0120] Suppose we need to generate semantic identification IDs for four types of data objects (Beauty, Sports, Toys, Instruments) disclosed by Amazon-ESCI. The specific flowchart is as follows Figure 3 as shown. Specifically:
[0121] S1. Select some data objects as training data and embed their object feature content. Here, we can use a general pre-trained text encoder (such as Sentence-T5 and BERT) to convert the text features of the data objects into semantic embeddings.
[0122] S2. Set hyperparameters, including the number and size of the codebooks, etc. Determine whether the hyperparameters conform to the tree structure. If they do, enter step S3; otherwise, wait for the user to input the correct hyperparameters. Examples of non-conforming cases are as follows: If the size of the root node codebook (i.e., the number of codes in the codebook) is set to 4, it means that each code in this codebook is expected to represent a data type (Beauty, Sports, Toys, or Instruments). Then there are 4 codebooks at the next level, and the user needs to set the sizes of the 4 codebooks respectively. If only 3 are set, it does not conform to the tree structure.
[0123] S3. Construct a semantic identification coding tree structure. Suppose the user sets the root node tree structure of the 0th layer in the tree structure to m = 4; there are m, that is, 4 codebooks in the 1st layer of the tree structure, and the size of each codebook can be set respectively as There are codebooks in the 2nd layer. If the sizes are set respectively as Specifically, if m = 4, The total number of codebooks in the second layer = 4 + 6 + 2 + 4 = 16. If the numbers of these 16 codebooks are set to 8, 8, 8, 8, 16, 16, 16, 16, 16, 16, 32, 64, N, N, N, N respectively, the overall structure is as follows Figure 4 as shown.
[0124] S4. Semantic Identification Coding Framework Codebook Initialization: Sample a batch of pre-training data and perform codebook initialization based on kmeans to avoid the problem of codebook collapse. Specifically, for example, use 10,000 samples to perform kmeans calculation of m cluster centers, and the obtained representation vectors of the four cluster centers are assigned to the 4 codebooks of the root node. Then, subtract the representation of the cluster center from the representation of each sample in each cluster to obtain the residual vector, and then perform kmeans clustering on the residual vectors respectively. For example, if 2,000 samples out of 10,000 samples are clustered onto the code with id = 0, then subtract the code with id = 0 in the root node codebook from the representations of the 2,000 samples to obtain 2,000 residual vectors. Then perform kmeans calculation of 4 cluster centers on these 2,000 residual vectors (because the size of the child node codebook corresponding to the code with id = 0 is 4, so the number of kmeans cluster centers is 4), and so on, to obtain the initial values of the codebooks at all levels from top to bottom.
[0125] S5. Model Training. Use the above loss function to train the model to obtain the final model parameters.
[0126]
[0127] S6. Obtain the semantic identifications of all data objects. For each data object, we generate information representations at each level based on the semantic identification calculation algorithm description in Table 1 to obtain the semantic identifications of all data objects. Note: As long as the number and size of the codebooks are large enough, the uniqueness of the semantic identifications of data objects can be basically guaranteed.
[0128] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.
[0129] As Figure 6 shown, the present invention also proposes a data object semantic identification generation device, which includes:
[0130] Initial module, obtain the target data x to be identified and construct a tree-shaped codebook model, where each node of the tree-shaped codebook model represents a codebook; set the level corresponding to the root node of the tree to level 0;
[0131] Coding module, extract the latent representation z of the target data through an encoder;
[0132] Quantization module, at the 0th layer of the tree-shaped codebook model, that is, the root node of the tree, the initial residual is r0 = z; map r0 to the nearest embedding in the node codebook To quantify r0, and the quantization result is an information representation value of the semantic identifier; is used as the current residual; the embedding is used as the current embedding; the information representation value is used as the representation value of the current 0th level;
[0133] The middle module processes the first layer of the tree-shaped codebook model, selects the child node of the code corresponding to the representation value obtained in the previous step in this layer as the current codebook, and maps the current residual r1 to the nearest embedding in the current codebook to quantify r1. And update the current residual r2 by subtracting the result of the current embedding from the updated current residual The quantization result will be another information representation value of the semantic identifier;
[0134] The loop module updates the current representation value with the other information representation value, and updates the current embedding with the embedding Update the current embedding, and repeat the execution of the middle module n times again until reaching the codebook where the leaf node is located in the tree-shaped codebook. Aggregate all the information representation values to obtain a tuple composed of all the information representation values as the semantic identifier of the target data.
[0135] The described data object semantic identifier generation device, which includes:
[0136] The decoding module adds the quantization vectors of the codebooks to which each information representation value belongs to obtain a summation result z ’ , and the decoder decodes the summation result z ’ to obtain x ’ , and constructs a loss function loss: where where sg represents the gradient stop operation, and r i is the updated current residual when the middle module is executed for the i-th time; the encoder, the decoder, and the tree-shaped codebook are trained through the loss function loss.
[0137] The described data object semantic identifier generation device, for the classification task, the loss function loss includes a classification loss Loss function loss:
[0138]
[0139] represents the cross-entropy loss of the k-th layer of the tree-shaped codebook; according to the category to which the target data belongs, obtain its semantic category id as the label, and according to the information representation value and the label of the k-th layer of the tree-shaped codebook, obtain the cross-entropy loss of the k-th layer;
[0140] The database building module constructs a database based on the target data with the semantic identifier, where the semantic identifier is used as the index of the target data, and the semantic identifier has the semantic category id to which the target data belongs;
[0141] The retrieval module obtains the data to be retrieved, obtains the semantic id of the data to be retrieved according to the semantics of the data to be retrieved, and retrieves the target data in the database that has the semantic id of the data to be retrieved as the retrieval result corresponding to the data to be retrieved. Semantic ids with the same prefix indicate similar objects.
[0142] As Figure 7 shown, in another embodiment, the present invention also proposes a first electronic device A, including the above-mentioned data object semantic identifier generation device.
[0143] As Figure 8 shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to acquire the target data x to be identified, such as video, picture, text, etc. The information display device D is used to display the semantic identifier analyzed by the present invention. The semantic identifier records the id number of the target data x belonging to a special category, which is equivalent to classifying the target data x.
[0144] Among them, the information display device D can process and organize the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be preset manually. For example, the data output by the first electronic device A is visually displayed, and it can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. The key information specified by the user is presented to the user, and the user can understand this information more timely without having to access the secondary page or scroll the page, saving the user's operation. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the key information of the user according to the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.
[0145] The present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the data object semantic identifier generation method provided by the above-mentioned various methods.
[0146] In another embodiment, the present invention further provides a storage medium VIII for storing a computer program for executing the method for generating semantic identification of data objects. It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM).
[0147] Figure 9 FIG. shows a schematic block diagram of a second electronic device 1000 that may be used to implement embodiments of the present invention. The second electronic device 1000 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.
[0148] The second electronic device 1000 includes a computing unit Ⅰ, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory Ⅱ (ROM) or computer programs loaded from a storage medium Ⅷ into a random access memory (RAM) Ⅲ. In the RAM Ⅲ, various programs and data required for the operation of the device 1000 can also be stored. The computing unit Ⅰ, the ROM Ⅱ, and the RAM Ⅲ are connected to each other via a bus Ⅳ. An input / output (I / O) interface Ⅴ is also connected to the bus Ⅳ.
[0149] Multiple components in the second electronic device 1000 are connected to the I / O interface Ⅴ, including: an input unit Ⅵ, such as a keyboard, a mouse, etc.; an output unit Ⅶ, such as various types of displays, speakers, etc.; a storage medium Ⅷ, such as a magnetic disk, an optical disc, etc.; and a communication unit Ⅸ, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit Ⅸ allows the second electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0150] The computing unit Ⅰ can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit Ⅰ include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit Ⅰ executes the various methods and processes described above, such as method steps S1 - S5. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage medium Ⅷ. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM Ⅱ and / or the communication unit Ⅸ. When the computer program is loaded into the RAM Ⅲ and executed by the computing unit Ⅰ, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit Ⅰ can be configured to execute the method in any other appropriate way (e.g., by means of firmware).
[0151] Although the embodiments of the present invention have been disclosed as above, it is not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the illustrations shown and described herein.
Claims
1. A method for generating semantic identifiers of data objects, characterized in that, Including: An initial step of obtaining target data x to be identified and constructing a tree-shaped codebook model, where each node of the tree-shaped codebook model represents a codebook; Setting the level corresponding to the root node of the tree as level 0; An encoding step of extracting a latent representation z of the target data through an encoder; Quantization step. At the 0th layer of the tree-shaped codebook model, which is the root node of the tree, the initial residual is r0 = z; map r0 to the nearest embedding in the node codebook to quantize r0, and the quantization result is an information representation value of the semantic identifier; Use it as the current residual; Use the embedding Taking the information representation value as the representation value of the current level 0; Intermediate step, process the first layer of the tree-shaped codebook model, select the child node of the code corresponding to the characterization value obtained in the previous step in this layer as the current codebook, and map the current residual r1 to the nearest embedding in this current codebook to quantify r1. And update the current residual with the result obtained by subtracting the current embedding from the updated current residual r2 The quantization result is another information characterization value with a semantic identifier; A loop step to update the current representation value with the other information representation value and with the embedding Update the current embedding, repeat the execution of this intermediate step n times again until reaching the codebook where the leaf node in the tree-shaped codebook is located, aggregate all the information representation values, and obtain a tuple composed of all the information representation values as the semantic identifier of the target data.
2. The method for generating a semantic identifier of a data object according to claim 1, wherein Including: Decoding step: Add the quantization vectors of the codebook to which each information representation value belongs to obtain a summation result z ’ , and the decoder decodes the summation result z ’ to obtain x ’ , construct a loss function loss: where where sg represents a gradient stop operation, and r i is the updated current residual at the i-th execution of this intermediate step; train the encoder, the decoder, and the tree-shaped codebook through this loss function loss 3. The method for generating a semantic identifier of a data object according to claim 1, wherein For the classification task, the loss function loss includes the classification loss Loss function loss: Represents the cross-entropy loss of the k-th layer of the tree-shaped codebook; according to the category to which the target data belongs, obtain its semantic category id as the label, and according to the information representation value and the label of the k-th layer of the tree-shaped codebook, obtain the cross-entropy loss of the k-th layer.
4. The method for generating a semantic identifier of a data object according to any one of claims 1 to 3, characterized in that, Including: A database building step of building a database based on the target data with the semantic identifier, where the semantic identifier is used as the index of the target data, and the semantic identifier has the semantic category id to which the target data belongs; A retrieval step of obtaining data to be retrieved, obtaining the semantic id of the data to be retrieved according to the semantics of the data to be retrieved, and retrieving the target data in the database that has the semantic id of the data to be retrieved as the retrieval result corresponding to the data to be retrieved. Semantic ids with the same prefix indicate similar objects.
5. A data object semantic identification generation device, characterized in that Including: An initial module of obtaining target data x to be identified and constructing a tree-shaped codebook model, where each node of the tree-shaped codebook model represents a codebook; Setting the level corresponding to the root node of the tree as level 0; An encoding module of extracting a latent representation z of the target data through an encoder; Quantization module, at the 0th layer of the tree-shaped codebook model, i.e., the root node of the tree, the initial residual is r0 = z; map r0 to the nearest embedding in the node codebook to quantize r0, and the quantization result is an information representation value of the semantic identifier; Take it as the current residual; Take the embedding as the current embedding; Taking the information representation value as the representation value of the current level 0; The intermediate module processes the first layer of the tree-shaped codebook model, selects the child node of the code corresponding to the feature value obtained in the previous step in this layer as the current codebook, and maps the current residual r1 to the nearest embedding in this current codebook. To quantize r1. And update the current residual with the result obtained by subtracting the current embedding from the updated current residual r2. The quantization result is another feature value of the semantic identifier; A loop module updates the current representation value with the other information representation value and embeds Update the current embedding, and repeat the execution of the intermediate module n times again until the codebook where the leaf node in the tree-shaped codebook is traversed. Aggregate all the information representation values to obtain a tuple composed of all the information representation values as the semantic identifier of the target data.
6. The data object semantic identification generation device according to claim 5, characterized in that, Including: The decoding module adds the quantization vectors of the codebook to which each information representation value belongs to obtain a summation result z ’ , and the decoder decodes the summation result z ’ to obtain x ’ , and constructs a loss function loss: where where sg represents a gradient stop operation, and r i is the current residual updated when the intermediate module is executed for the i-th time; the encoder, the decoder, and the tree-shaped codebook are trained through the loss function loss.
7. The data object semantic identification generation device according to claim 5 or 6, characterized in that, For the classification task, the loss function loss includes the classification loss Loss function loss: Represent the cross-entropy loss of the k-th layer of the tree-shaped codebook; according to the category to which the target data belongs, obtain its semantic category id as a label, and based on the information representation value and the label of the k-th layer of the tree-shaped codebook, obtain the cross-entropy loss of the k-th layer; A database building module of building a database based on the target data with the semantic identifier, where the semantic identifier is used as the index of the target data, and the semantic identifier has the semantic category id to which the target data belongs; A retrieval module of obtaining data to be retrieved, obtaining the semantic id of the data to be retrieved according to the semantics of the data to be retrieved, and retrieving the target data in the database that has the semantic id of the data to be retrieved as the retrieval result corresponding to the data to be retrieved. Semantic ids with the same prefix indicate similar objects.
8. An electronic device, characterized in that, Including a data object semantic identifier generation device according to claims 5-7, where the electronic device is connected to an information display device, and the information display device is used to display the semantic identifier of the target data with display parameters, attributes set by the user, or through an artificial intelligence model.
9. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the data object semantic identifier generation method according to any one of claims 1-4.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the data object semantic identifier generation method according to any one of claims 1-4.