Sparse auto-encoder-based lexical element construction method and system
Through the word unit construction method of sparse autoencoders, the problems of training imbalance and embedding collapse in large language model recommendation systems are solved. Through sparse representation and orthogonal constraint optimization, the quality of item representation and recommendation performance are improved.
Patent Information
- Application Number
- CN202511171124.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing methods suffer from training imbalance and embedding collapse problems in recommendation systems based on large language models, which results in the model being unable to fully utilize the codebook vector, affecting the recommendation performance.
The word unit construction method of sparse autoencoder is adopted. The similarity between the codebook vector and the original embedding is calculated through the weight sharing strategy. Top-K sparsification is performed to generate sparse representations. The semantic embedding is reconstructed through the decoder. The model is optimized by combining joint reconstruction loss, orthogonal constraint and diversity regularization.
It significantly improves the quality of item representation and recommendation performance, alleviates the problems of training imbalance and embedding collapse, and improves the representation ability and generalization performance of the model.
Smart Images

Figure CN120725010A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of recommendation systems, and more specifically to a word unit construction method and system based on a sparse autoencoder. Background Art
[0002] With the widespread application of large-scale language models (LLMs) in recommendation systems, using LLMs for recommendations has become a new paradigm. However, in generative recommendation systems based on pre-trained large language models, the model input and output are typically presented in the form of tokens. Therefore, converting user interaction sequences into discrete token sequences is a key step in data processing and directly affects the performance of sequence recommendation tasks. Traditional methods assign a unique ID token to each item as its discrete representation. However, as the item set gradually increases, this method requires adding all item IDs as new tokens to the pre-trained model's vocabulary, affecting the performance of the original model, increasing computational overhead, and limiting scalability. Some work has also attempted to directly use item titles as indexes for the input and output of large models. However, this approach fails to fully utilize other information about the item, and the large model generation space is huge, so the results may not be mapped to real items.
[0003] Although existing methods have solved the above problems to some extent, they still have two major limitations: 1. Training imbalance problem. Existing vector quantization methods usually rely on multiple codebooks and perform hierarchical quantization of the semantic features of items through residual connections. However, this design suffers from a serious training imbalance problem during training. Specifically, each codebook usually only maps the input feature to its nearest single codebook vector, resulting in most codebook vectors not being activated throughout the training process. For example, two different items may always be mapped to the same codebook vector, while other codebook vectors remain untrained. This unbalanced update distribution prevents the model from fully utilizing the entire codebook, resulting in inefficient training and limiting the model's representation capabilities.
[0004] 2. Embedding collapse problem. Codebook vectors in existing vector quantization methods often exhibit high similarity and dependence, leading to the degradation of the representation space (Embedding Collapse). From a geometric perspective, this degradation manifests as linear dependence of the codebook vectors, making the dimension of the representation space far below the theoretical upper limit. For example, in a three-dimensional space, the code vector may only cover a two-dimensional subspace, significantly reducing the model's representation ability and generalization performance. This embedding collapse problem not only limits the model's ability to capture the semantic features of items, but also makes the generated discrete semantic IDs unable to effectively express the semantic similarity between items. Summary of the Invention
[0005] In this embodiment, a word unit construction method, system, electronic device and storage medium based on a sparse autoencoder are provided to solve the embedding collapse problem in related technologies, which not only limits the model's ability to capture the semantic features of objects, but also makes the generated discrete semantic ID unable to effectively express the semantic similarity between objects.
[0006] In a first aspect, an embodiment of the present invention provides a word unit construction method based on a sparse autoencoder, the word unit construction method based on a sparse autoencoder comprising the following steps: Get the semantic embedding vector of the item; A weight sharing strategy is adopted to calculate the similarity between the codebook vector and the original embedding through the weight-sharing codebook; Perform top-K sparsification on the similarity, retain the K codebook vectors with the largest similarity values, and set the rest to zero to generate a sparse representation; Generate a discrete word sequence based on the position and value of non-zero elements in the sparse representation; Reconstruct semantic embedding through the decoder; Joint reconstruction loss, orthogonality constraint, and diversity regularization optimization model.
[0007] In an optional embodiment, the weight sharing strategy is: Make the encoder and decoder share a trainable codebook , where N represents the size of the codebook, d is the dimension of the semantic embedding, Codebook The i-th codebook vector in ; The encoding process is regarded as calculating the similarity between each codebook vector in the codebook and the original embedding, and the decoding process is regarded as reconstructing the original embedding based on the highly similar codebook vectors.
[0008] In an optional embodiment, the similarity between the codebook vector and the original embedding is calculated using a weight-shared codebook, including: For a given embedding representation, its latent space representation is calculated through the encoder: in, is the semantic embedding of the item, The codebook matrix is multiplied by the semantic embedding vector, indicating that all codebook vectors in the codebook are calculated as the inner product with the semantic embedding of the item as the similarity. is the first codebook vectors.
[0009] In an optional embodiment, top-K sparsification is performed on the similarity, retaining the K codebook vectors with the largest similarity values and setting the rest to zero, to generate a sparse representation, including: Keep the K codebook vectors that are most relevant to the input embedding representation, and set the rest to zero to obtain a sparse representation ; ; in, is the value of the jth element in the latent space representation of the i-th item, represents the subscripts of all non-zero elements in the sparse representation of the i-th item, that is, the subscripts corresponding to the top-K largest elements. The top-K largest elements refer to the non-zero positions corresponding to the K codebook vectors with the highest similarity scores.
[0010] In an optional embodiment, the discrete word sequence is generated in the following manner: Sort the non-zero elements in the sparse representation from large to small, and convert the corresponding codebook index into a word identifier.
[0011] In an optional embodiment, reconstructing the semantic embedding by the decoder includes: The original embedding is reconstructed by weighted combination of different codebook vectors in a shared weighted manner during the decoding phase. ; in, is the output of the model, is the value of the jth element in the latent space representation of the i-th item, Represents the subscript corresponding to the top-K largest elements in the sparse representation of the i-th item, where the top-K largest elements refer to the non-zero positions corresponding to the K codebook vectors with the highest similarity scores. represents the jth codebook vector in the codebook.
[0012] In an optional embodiment, the joint reconstruction loss, orthogonality constraint, and diversity regularization optimization model includes: Reconstruction loss: ; in is the original semantic embedding of the input, The reconstructed semantic embedding output by the model; Orthogonality constraints: , ; in, is the size of the batch data, represents a one-hot operation, Indicates that only vectors are retained The value of the L-th largest element in the array, and the rest of the elements are set to zero. express The result after the one-hot operation; Used to calculate the cosine similarity of two items i and j in the same batch at the Lth mark position; Diversity Regularization Optimization: .
[0013] in represents the identity matrix, represents the codebook matrix.
[0014] Compared with the prior art, the word unit construction method based on sparse autoencoder of the present invention has the following beneficial effects: This method significantly improves the quality of item representation and recommendation performance through sparse representation and orthogonal constraint optimization. Specifically, it utilizes a sparse autoencoder combined with a trainable codebook to directly learn a sparse representation of item semantic features. Orthogonal constraints are used to ensure the independence and semantic uniqueness of the codebook vectors, effectively alleviating training imbalance and embedding collapse. Experimental results demonstrate that this method outperforms existing methods on multiple benchmark datasets, significantly improving the quality of item representation and the performance of the recommendation system.
[0015] In a second aspect, an embodiment of the present invention provides a word unit construction system based on a sparse autoencoder, comprising: Acquisition module, obtains the semantic embedding vector of the item; The similarity calculation module adopts a weight sharing strategy to calculate the similarity between the codebook vector and the original embedding through the weight-sharing codebook; The sparse representation generation module performs top-K sparsification on the similarity, retaining the K codebook vectors with the largest similarity values and setting the rest to zero to generate a sparse representation; The discrete word sequence generation module generates a discrete word sequence based on the position and value of non-zero elements in the sparse representation; The reconstruction module reconstructs the semantic embedding through the decoder; Joint optimization module, which jointly optimizes the model with reconstruction loss, orthogonality constraint and diversity regularization.
[0016] In a third aspect, an embodiment of the present invention provides an electronic device comprising a processor, a communication interface, a memory and a bus, wherein the processor, the communication interface and the memory communicate with each other through the bus, and the processor can call logic instructions in the memory to execute the steps of the method provided in the first aspect.
[0017] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the word unit construction method based on the sparse autoencoder as described in the first aspect.
[0018] Compared with the prior art, the beneficial effects of the word unit construction system, electronic device and storage medium based on sparse autoencoder of the present invention are the same as those of the word unit construction method based on sparse autoencoder described in the first aspect, so they will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 Flowchart of a word unit construction method based on a sparse autoencoder in an embodiment of the present invention; Figure 2 This is a flow chart of model training in an embodiment of the present invention; Figure 3 1 is a structural block diagram of a word unit construction system based on a sparse autoencoder in an embodiment of the present invention; Figure 4 2 is a structural block diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to more clearly understand the purpose, technical solutions and advantages of this application, this application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0022] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meanings as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "the," "these," and similar expressions in this application do not denote limitations on quantity and may be singular or plural. The terms "comprise," "include," "have," and any variations thereof, as used in this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include unlisted steps or modules (units) or other steps or modules (units) inherent to the process, method, product, or device. The terms "connected," "connected," "coupled," and similar expressions used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used in this application, "plurality" means two or more. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone; A and B exist simultaneously; or B exists alone. Generally, the character " / " indicates that the objects in the preceding and following relationship are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0023] In recommendation systems based on large-scale language models (LLMs), obtaining discrete tokens for items is a key step in sequential recommendation tasks. However, existing vector quantization (VQ) methods have significant shortcomings in dealing with training imbalance and embedding collapse. To overcome these challenges, this paper proposes a token construction framework (RecTok) based on a sparse autoencoder (SAE), which can adaptively select and learn important semantic features and discretely represent items. We further propose a new optimization objective to address the problems of representation diversity and embedding collapse. Our core idea is to directly learn sparse representations of items through a sparse autoencoder and introduce orthogonality constraints and diversity regularization to optimize the representation space of the codebook. Next, we will introduce our RecTok method in detail.
[0024] In an embodiment of the present invention, a word unit construction method based on a sparse autoencoder is provided, which is applied to a recommendation system based on a large-scale language model. Figure 1 This is a flow chart of the word unit construction method based on sparse autoencoder of the present invention, as shown in FIG. Figure 1 As shown, the process includes the following steps: S100, obtaining the semantic embedding vector of the item; S200, using a weight sharing strategy, and calculating the similarity between the codebook vector and the original embedding through the weight-sharing codebook; In this embodiment, the weight sharing strategy is: Make the encoder and decoder share a trainable codebook , where N represents the size of the codebook and d is the dimension of the semantic embedding; The encoding process is regarded as calculating the similarity between each codebook vector in the codebook and the original embedding, and the decoding process is regarded as reconstructing the original embedding based on the highly similar codebook vectors.
[0025] Specifically, this solution uses a sparse autoencoder (SAE) to map continuous semantic embeddings to discrete code vectors. To simplify the model and ensure consistency in the semantic space, this solution adopts a weight sharing strategy. More specifically, the encoder and decoder share a trainable codebook. The encoding process can be viewed as calculating the similarity between each codebook vector and the original embedding, and the decoding process can be viewed as reconstructing the original embedding based on highly similar codebook vectors. This design not only reduces model complexity but also ensures consistency between the encoding and decoding processes in the semantic space.
[0026] The similarity between the codebook vector and the original embedding is calculated through the weight-sharing codebook, including: For a given embedding representation, its latent space representation is calculated through the encoder: in, is the semantic embedding of the item, The codebook matrix is multiplied by the semantic embedding vector, indicating that all codebook vectors in the codebook are calculated as the inner product with the semantic embedding of the item as the similarity. is the first codebook vectors.
[0027] S300, performing top-K sparsification on the similarity, retaining the K codebook vectors with the largest similarity values, and setting the rest to zero, to generate a sparse representation; Perform top-K sparsification on the similarity, retain the K codebook vectors with the largest similarity values, and set the rest to zero to generate a sparse representation, including: Keep the K codebook vectors that are most relevant to the input embedding representation, and set the rest to zero to obtain a sparse representation ; ; ; in, is the value of the jth element in the latent space representation of the i-th item, Represents the subscripts of all non-zero elements in the sparse representation of the i-th item, that is, the subscripts corresponding to the top-K largest elements, where the top-K largest elements refer to the non-zero positions corresponding to the K codebook vectors with the highest similarity scores.
[0028] S400, generating a discrete word sequence according to the position and numerical value of non-zero elements in the sparse representation; In this embodiment, the discrete word sequence is generated as follows: Sort the non-zero elements in the sparse representation from large to small, and convert the corresponding codebook index into a word identifier.
[0029] Specifically, given an embedding representation , the encoder calculates its latent space representation. In order to ensure sparsity, the top-K sparsification mechanism is introduced to retain the K codebook vectors most relevant to the input embedding, and the rest are set to zero to obtain a sparse representation , the above sparse coding strategy makes It can be represented by the K most relevant codebook vectors, and the information contained in it is concentrated on the most relevant K codebook vectors, and adaptively selected and updated in the subsequent optimization process. , a discrete ID representation is obtained based on the subscript of each position and the size of the element. For example, the sparse representation of an item , where the element at the third position is the largest, followed by the element at the first position, so the discrete ID of the item is<T_3> 、<T_1> .
[0030] S500, reconstructing semantic embedding through decoder; The semantic embedding is reconstructed through the decoder, including: The original embedding is reconstructed in the decoding stage by weighted combination of different codebook vectors in a shared weight manner.
[0031] The decoder weights are the matrix of all codebook vectors, which is the transpose of the encoder weights. By sharing weights, the original embedding is reconstructed during the decoding phase by weighted combination of different codebook vectors. This allows us to retain the most relevant K codebook vectors selected during the encoding phase based on semantic similarity and reconstruct the original input: ; in, is the output of the model, is the value of the jth element in the latent space representation of the i-th item, Indicates the subscript corresponding to the top-K largest elements in the sparse representation of the i-th item, represents the jth codebook vector in the codebook.
[0032] S600, joint reconstruction loss, orthogonality constraint, and diversity regularization optimization model.
[0033] In this solution, the optimization objective of the model consists of a reconstruction objective, an orthogonal constraint, and a diversity regularization. The reconstruction objective focuses on the consistency of the sparse representation after encoding by the sparse autoencoder with the original output, and calculates the mean square error between the decoder output and the original input: ; in is the original semantic embedding of the input, The reconstructed semantic embedding output by the model.
[0034] The diversity constraint minimizes the cosine similarity of the Lth token position of items in the same batch, promoting balanced utilization of the codebook while preserving the semantic similarity of the first L-1 tokens: ; ; in, is the size of the batch data, represents a one-hot operation, Indicates that only vectors are retained The value of the L-th largest element in the array, and the rest of the elements are set to zero. express The result after the one-hot operation; Used to calculate the cosine similarity of two items i and j in the same batch at the Lth mark position; In addition, to ensure the integrity of the codebook semantic space, the similarity and dependency between codebook vectors in the codebook should be reduced. Therefore, the orthogonal constraint is used to further optimize the codebook expression capability: ; in, represents the identity matrix, represents the codebook matrix.
[0035] Overall, the RecTok method is based on a simple and effective sparse autoencoder, which converts the sparse representation of semantics into discrete IDs for subsequent LLM training. This method is not only computationally efficient but also scalable to large-scale datasets.
[0036] The RecTok method proposed in this study demonstrates significant advantages and positive results in the task of discrete item representation in recommendation systems based on large language models. First, RecTok integrates a sparse autoencoder with a codebook structure, sparsely activates and optimizes the K most relevant codebook vectors. It also introduces orthogonal constraints and diversity regularization to optimize the codebook's expressiveness, effectively overcoming the training imbalance and embedding collapse issues of traditional word-unit construction methods based on vector quantization. Compared with existing methods, this scheme not only enhances the ability to capture item semantic features, but also reduces computational complexity through single-codebook sparse coding, achieving efficient alignment between the recommendation task and the autoregressive properties of LLMs.
[0037] The experimental results are shown in Table 1. RecTok comprehensively outperforms various sequence recommendation baselines on three datasets of Amazon's e-commerce platform (Musical Instruments, Video Games, and Baby Products). On the Games dataset, RecTok's HR@1 reached 0.0223 and HR@5 reached 0.0637, which are 6.70% and 3.74% higher than LC-Rec, respectively. On the Instruments dataset, HR@5 and NDCG@5 increased by 19.4% and 17.3% compared to the variant without sparse activation, verifying the key role of sparse representation in semantic modeling. Through the collaborative design of sparse coding and geometric constraints, this method not only retains the fine-grained representation of item semantics, but also adapts to the autoregressive generation paradigm of LLM, showing significant engineering application potential in large-scale dynamic scenarios such as e-commerce and content recommendation. Table 1: CVR prediction results of different methods on the Criteo dataset
[0038] For example, Figure 2 As shown, Figure 2 The following is a flowchart showing how to train the model and generate recommendation results. Specifically, the process is as follows: Data collection and preprocessing: Collect all items in the dataset and their corresponding text information such as title, category, description, etc. Encode the item text information through the pre-trained text encoding model to obtain semantic embedding.
[0039] Model initialization: The codebook vector is initialized using a uniform distribution, and the encoder and decoder weights are tied and transposed to each other.
[0040] Model parameter update: The sparse autoencoder is optimized through multi-objective joint optimization, which is achieved through optimization algorithms (such as stochastic gradient descent) to balance computational efficiency and effectiveness.
[0041] Get the discrete representation of the item: The trained model is used to infer the semantic embeddings of all items and assign discrete tokens based on the obtained sparse activation results.
[0042] Fine-tuning the language model: According to the mapping relationship between items and tokens in the previous step, the user's interaction sequence is represented by tokens and a language model prompt word fine-tuning model is constructed.
[0043] Generate recommendation results: After training is completed, the next item is predicted in the form of word units, and finally the specific item ID or title is mapped by the reverse process of step 4.
[0044] The RecTok method proposed in this proposal can be widely applied to language model-based sequence recommendation tasks in scenarios such as e-commerce recommendations, content platforms, and streaming services. In practical applications, RecTok effectively improves item tagging quality and recommendation accuracy in various large-scale dynamic scenarios. For example, on large e-commerce platforms, merchants can use RecTok to convert product metadata (such as titles and descriptions) into discrete tag sequences using a sparse autoencoder and integrate them into the LLM recommendation generation process. This method uses top-K sparsification to retain the most relevant code vectors and incorporates orthogonal constraints to prevent semantic collapse. This enables the LLM to more accurately capture the semantic associations between products in a user's historical interaction sequences, thereby optimizing the conversion rate of recommendation modules such as "Guess You Like" (likely referring to a specific topic or feature). In video streaming scenarios, RecTok can convert textual metadata (such as plot synopses and tags) of film and television content into semantically compact tag sequences, helping the LLM model the temporal patterns of a user's viewing history and dynamically adjust its recommendation strategy to adapt to real-time preferences. For example, when a user continuously watches science fiction films, the tag sequences generated by RecTok can enhance the LLM's ability to capture the "science fiction" semantic cluster and prioritize recommendations for content within the same category.
[0045] Similarly, on news information platforms, RecTok can semantically encode the titles and summaries of news articles. The generated discrete tags can be effectively integrated into the LLM's autoregressive recommendation model, addressing the problem that traditional ID taggers cannot capture the timeliness of content. By encouraging balanced utilization of the codebook through diversity regularization, RecTok enables the LLM to quickly update the tag representations of relevant topics when processing breaking news events, achieving real-time iteration of recommended content. Furthermore, in the game app store recommendation scenario, RecTok can convert metadata such as game genre and gameplay description into structured tags, helping the LLM model user game preference sequences. Especially when processing newly released games, sparse coding and codebook optimization can accelerate the semantic integration of new content and reduce cold-start recommendation errors. These application examples demonstrate that RecTok, through the deep coupling of semantic tagging with the autoregressive characteristics of the LLM, has demonstrated flexible adaptability and engineering practicality in large-scale recommendation systems across various domains.
[0046] The embodiment of the present invention also provides a word unit construction system based on a sparse autoencoder, which is used to implement the above-mentioned method embodiment, and will not be repeated hereafter. The terms "module", "unit", "sub-unit", etc. used below can implement a combination of software and / or hardware for predetermined functions. Although the system described in the following embodiments is preferably implemented in software, implementation by hardware or a combination of software and hardware is also possible and conceivable.
[0047] like Figure 3 As shown, Figure 3 : is a structural block diagram of a word unit construction system based on a sparse autoencoder in the present invention, which includes: Acquisition module 101, acquiring the semantic embedding vector of the item; Similarity calculation module 102, adopting a weight sharing strategy, calculates the similarity between the codebook vector and the original embedding through the weight-sharing codebook; The sparse representation generation module 103 performs top-K sparsification on the similarity, retains the K codebook vectors with the largest similarity values, and sets the rest to zero to generate a sparse representation; A discrete word unit sequence generation module 104 generates a discrete word unit sequence according to the position and value of non-zero elements in the sparse representation; Reconstruction module 105, reconstructs semantic embedding through decoder; The joint optimization module 106 jointly optimizes the model using reconstruction loss, orthogonality constraint, and diversity regularization.
[0048] Figure 4 A structural block diagram of an electronic device provided by an embodiment of the present invention, such as Figure 4As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the following method: S100, obtaining the semantic embedding vector of the item; S200, using a weight sharing strategy, and calculating the similarity between the codebook vector and the original embedding through the weight-sharing codebook; S300, performing top-K sparsification on the similarity, retaining the K codebook vectors with the largest similarity values, and setting the rest to zero, to generate a sparse representation; S400, generating a discrete word sequence according to the position and numerical value of non-zero elements in the sparse representation; S500, reconstructing semantic embedding through decoder; S600, joint reconstruction loss, orthogonality constraint, and diversity regularization optimization model.
[0049] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0050] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method provided in the above embodiments is implemented.
[0051] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or certain parts of the embodiment.
[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A word unit construction method based on sparse autoencoders, applied to recommendation systems based on large-scale language models, characterized by: The steps include: Get the semantic embedding vector of the item; A weight sharing strategy is adopted to calculate the similarity between the codebook vector and the original embedding through the weight-sharing codebook; Perform top-K sparsification on the similarity, retain the K codebook vectors with the largest similarity values, and set the rest to zero to generate a sparse representation; Generate a discrete word sequence based on the position and value of non-zero elements in the sparse representation; Reconstruct semantic embedding through the decoder; Joint reconstruction loss, orthogonality constraint, and diversity regularization optimization model.
2. The word unit construction method based on sparse autoencoder according to claim 1 is characterized in that The weight sharing strategy is: Make the encoder and decoder share a trainable codebook , where N represents the size of the codebook, d is the dimension of the semantic embedding, Codebook The i-th codebook vector in ; The encoding process is regarded as calculating the similarity between each codebook vector in the codebook and the original embedding, and the decoding process is regarded as reconstructing the original embedding based on the highly similar codebook vectors.
3. The word unit construction method based on sparse autoencoder according to claim 2, characterized in that The similarity between the codebook vector and the original embedding is calculated through the weight-sharing codebook, including: For a given embedding representation, its latent space representation is calculated through the encoder: in, is the semantic embedding of the item, The codebook matrix is multiplied by the semantic embedding vector, indicating that all codebook vectors in the codebook are calculated as the inner product with the semantic embedding of the item as the similarity. is the first codebook vectors.
4. The word unit construction method based on sparse autoencoder according to claim 1, characterized in that Perform top-K sparsification on the similarity, retain the K codebook vectors with the largest similarity values, and set the rest to zero to generate a sparse representation, including: Keep the K codebook vectors that are most relevant to the input embedding representation, and set the rest to zero to obtain a sparse representation ; ; in, is the value of the jth element in the latent space representation of the i-th item, Represents the subscript corresponding to the top-K largest elements in the sparse representation of the i-th item, where the top-K largest elements refer to the non-zero positions corresponding to the K codebook vectors with the highest similarity scores.
5. The word unit construction method based on sparse autoencoder according to claim 1 is characterized in that The discrete word sequence generation method is: Sort the non-zero elements in the sparse representation from large to small, and convert the corresponding codebook index into a word identifier.
6. The word unit construction method based on sparse autoencoder according to claim 1, characterized in that: The semantic embedding is reconstructed through the decoder, including: The original embedding is reconstructed by weighted combination of different codebook vectors in a shared weighted manner during the decoding phase. ; in, is the output of the model, is the value of the jth element in the latent space representation of the i-th item, Indicates the subscript corresponding to the top-K largest elements in the sparse representation of the i-th item. The top-K largest elements refer to the non-zero positions corresponding to the K codebook vectors with the highest similarity scores. represents the jth codebook vector in the codebook.
7. The word unit construction method based on sparse autoencoder according to claim 1, characterized in that: The joint reconstruction loss, orthogonality constraint, and diversity regularization optimization model includes: Reconstruction loss: ; in, is the original semantic embedding of the input, The reconstructed semantic embedding output by the model; Orthogonality constraints: , ; in, is the size of the batch data, represents a one-hot operation, Indicates that only vectors are retained The value of the L-th largest element in the array, and the rest of the elements are set to zero. express The result after the one-hot operation; Used to calculate the cosine similarity of two items i and j in the same batch at the Lth mark position; Diversity Regularization Optimization: ; in represents the identity matrix, represents the codebook matrix.
8. A word unit construction system based on sparse autoencoder, characterized in that: include: Acquisition module, obtains the semantic embedding vector of the item; The similarity calculation module adopts a weight sharing strategy to calculate the similarity between the codebook vector and the original embedding through the weight-sharing codebook; The sparse representation generation module performs top-K sparsification on the similarity, retaining the K codebook vectors with the largest similarity values and setting the rest to zero to generate a sparse representation; The discrete word sequence generation module generates a discrete word sequence based on the position and value of non-zero elements in the sparse representation; The reconstruction module reconstructs the semantic embedding through the decoder; Joint optimization module, which jointly optimizes the model with reconstruction loss, orthogonality constraint and diversity regularization.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the word unit construction method based on the sparse autoencoder according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the word unit construction method based on the sparse autoencoder are implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
System and method for realizing handwriting identification based on sparse auto-encoding codebook
CN106529490A
Discretization processing-based model training method, prediction method and device
CN118349897A
Codebook compression for vector quantized neural networks
US20250245494A1
Drug property prediction method and device based on graph neural network using vector quantization
WO2024250185A1
Cited By
Feature editing method for large model content security
CN121212369A
A feature editing method for large model content security
CN121212369B