Multi-modal sequence recommendation method based on KAN network structure
By adopting the KAN network structure method in the multimodal sequence recommendation system, a feature interaction module within and between multimodals is constructed, which solves the problems of high computational complexity and difficulty in mining user preference information in the existing system, and achieves efficient and accurate user preference mining and recommendation effects.
Patent Information
- Application Number
- CN202510534576.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing multimodal sequence recommendation system has the problems of high time calculation complexity, low space utilization efficiency, and it is difficult to effectively explore user preference information under various modes.
Using a multimodal sequence recommendation method based on KAN network structure, by constructing the Token-Channel serial single-modal KAN feature interaction module and the multimodal KAN feature interaction fusion module, explicitly model the information correlation between multi-source multimodal heterogeneous data, deeply explore the user's fine-grained interest preferences under different modes, and perform multimodal feature cascade operations to generate the final embedded learning representation.
It significantly reduces the computational complexity, improves the system operation efficiency, can more accurately meet the potential needs of users, and deeply explores the behavioral preference characteristics of users under multimodal conditions.
Smart Images

Figure CN120067460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and in particular to a multi-modal sequence recommendation method based on a KAN network structure. Background Art
[0002] With the booming development of Internet multimedia sharing platforms, recommendation systems have become an indispensable core component. The recommendation system aims to mine the potential preference intention information of users based on the user's historical behavior records, and provide personalized recommendation services for different users in an implicit way. Compared with traditional recommendation algorithms based on the id paradigm, the research on multi-modal recommendation algorithms is still in its infancy. The multi-modal recommendation algorithm mines the potential user preference information of users in different modalities by according to the item sequence and the multi-modal attribute features of the items that the user has interacted with historically.
[0003] However, the existing multi-modal sequence recommendation system frameworks generally have problems such as high time computational complexity, low space utilization efficiency, and difficulty in effectively mining user preference information in each modality. Summary of the Invention
[0004] To solve the above problems, the present invention provides a multi-modal sequence recommendation method based on a KAN network structure, which explicitly models the information association between multi-source multi-modal heterogeneous data, deeply mines the fine-grained interest preferences of users in different modalities, fully integrates multi-modal interest representations, and more accurately meets the potential needs of users. At the same time, the framework adopts a lightweight KAN network structure design, significantly reducing the computational complexity and improving the system operation efficiency, and has broad practical application value.
[0005] To achieve the above object, the present invention provides a multi-modal sequence recommendation method based on a KAN network structure, specifically including the following steps: Step S1: Given the user's historical interaction multi-modal item sequence as input, extract the corresponding modal features of the item sequence; Step S2: Construct a Token-Channel serial single-modal intra-KAN feature interaction module for the interaction between the new user and the item, optimize the modeling of the user's fine-grained interest preference for a single specific modality, and perform interaction learning on the single-modal representation of the item sequence; Step S3: Construct a Token-Channel serial multi-modal inter-KAN feature interaction and fusion module for the interaction between the new user and the item, fully integrate multi-modal feature learning, perform multi-modal feature concatenation operations on each single-modal feature matrix, perform item feature interaction and fusion respectively in the token dimension and the channel dimension, model and learn cross-modal user preference information, and generate the final embedding learning representation; Step S4: For the embedded learning representation generated in step S3, use the item prediction layer prediction module to predict the specific item to interact with the user next. Obtain a lower-dimensional representation through a multi-layer perceptron and use a classifier to screen and subdivide the item, obtaining the recommendation result.
[0006] Preferably, in step S2, constructing a new user-item interaction token-word list serial single-modal intra-KAN feature interaction module specifically includes the following steps: S21: Introduce Token-level and Channel-level KAN neural network modules, enhancing the expression ability and mutual understanding ability of modal information in the high-dimensional feature space; S22: Adopt an improved feature normalization mechanism to achieve feature consistency within a single modality and between cross-modalities through LayerNorm combined with TokenKAN and ChannelKAN modules; For the visual modality feature sequence 、text modality feature sequence 、modal feature sequence of identifier id , the function calculation formulas within each single modality are as follows: ; ; ; ; ; ; Among them, represents the visual modality information after Token-level feature interaction of the visual modality feature, represents the text modality information after Token-level feature interaction of the text modality feature, represents the identifier modality information after Token-level feature interaction of the identifier modality feature, and i represents information interaction in the column direction of the feature interaction matrix. represents the visual modality feature after Channel-level feature fusion of the visual modality information, represents the text modality feature after Channel-level feature fusion of the text modality information, represents the identifier modality feature after Channel-level feature fusion of the identifier modality information, and j represents information interaction in the row direction of the feature interaction matrix.
[0007] Preferably, in step S3, construct a new Token-Channel serial multi-modal inter-KAN feature interaction and fusion module for user-item interaction, specifically including the following steps: S31: Introduce the feature output generated by the single-modal intra-feature interaction module based on the KAN network in step S2 and perform a multi-modal feature concatenation operation on each single-modal feature matrix to obtain a multi-modal feature interaction matrix that needs to be fused , and the function formula is as follows: ; where represents a linear mapping dimensionality reduction module that reduces the features of to ; represents a concatenation operation that concatenates multi-source multi-modal feature matrices; S32: For the multi-modal feature matrix after concatenation, use normalization in the marked dimension to limit the feature value size to prevent gradient disappearance and gradient explosion during the training process; S33: Use the KAN neural network fusion interaction operation between multi-modalities to fuse the feature information in the marked dimension of multi-modal information, and use residual connection to inject the feature results before calculation to generate the intermediate step feature embedding representation result after fusion between multi-modalities , and the specific function formula is as follows: ; In the formula, represents layer normalization operation to prevent gradient disappearance and gradient explosion during the model training process, represents the multi-modal feature information after Token-level feature interaction, and i represents information interaction in the column direction of the feature interaction matrix.
[0008] S34: Based on the intermediate step feature embedding representation result after multi-modal fusion , perform normalization operation in the channel dimension and execute the KAN neural network interaction fusion operation between multi-modalities in the channel dimension to perform inter-modal feature interaction learning and generate the final feature embedding representation result , and the function formula is as follows: ; In the formula, represents the multi-modal feature information after Channel-level feature interaction, and j represents information interaction in the column direction of the feature interaction matrix.
[0009] Preferably, in step S31, a feature fusion scheme based on dual optimization of Token dimension and Channel dimension is proposed for complex feature interactions in multimodal data, and a multimodal feature interaction matrix of items is constructed to achieve efficient collaboration and information flow among multimodalities. The feature interaction weights among modalities are dynamically adjusted through distributed weight learning of the KAN network.
[0010] Preferably, in step S32, a feature optimization strategy for inter-modality difference alignment is used to achieve effective feature fusion and denoising within and between modalities through a layer-by-layer interaction mechanism, and the inter-modality differences include vision, text, and identifier sequences.
[0011] Preferably, in step S34, a progressive recommendation result generation mechanism is adopted, and L2 regularization and a layer-by-layer feature fusion strategy of a KAN network are introduced to improve the diversity and accuracy of the final recommendation results.
[0012] Therefore, the present invention adopts the above-mentioned multimodal sequence recommendation method based on the KAN network structure, which has the following beneficial effects: (1) The present invention introduces the KAN neural network as the basic computing unit, optimizes the neural network structure of the traditional attention computing mechanism based on quadratic time computational complexity, enhances the discriminative ability of representation, and improves the efficiency of the model.
[0013] (2) The present invention designs a novel Token-Channel serial multi-modal KAN feature interaction fusion module for user-item interaction, which intends to use the KAN network to interactively fuse the information within a specific single modality in each modality of the user interaction item sequence from two perspectives, namely the token dimension and the channel dimension, and to mine the user's fine-grained potential preference intention information in each modality.
[0014] (3) The present invention designs a novel Token-Channel serial multimodal KAN feature interaction fusion module for user-item interaction, and then performs feature interaction and fusion in the token dimension and channel dimension to effectively model cross-modal user preference information, generate the final embedding representation, and deeply mine the user's behavioral preference characteristics under multimodal conditions.
[0015] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of a multimodal sequence recommendation method based on a KAN network structure of the present invention; Figure 2 This is a detailed structural diagram of the feature interaction fusion module based on the KAN network structure in the present invention. Detailed implementation manners
[0017] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0018] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs.
[0019] The terms such as "including" or "comprising" used in the present invention mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements. The orientation or positional relationship indicated by terms such as "inside", "outside", "above", "below", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation to the present invention. When the absolute position of the described object changes, the relative positional relationship may also change accordingly. In the present invention, unless otherwise clearly defined and limited, terms such as "attaching" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected, or indirectly connected through an intermediate medium. It can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0020] A multi-modal sequence recommendation method based on the KAN network structure specifically includes the following steps: Step S1: Given the user's historical interaction multi-modal item sequence as input, extract the corresponding modal features of the item sequence; Step S2: Construct a novel Token-Channel serial single-modal intra-KAN feature interaction module for user-item interaction, and optimize the modeling of the user's fine-grained interest preferences under a single specific modality and the interaction learning of the single-modal intra-representation of the item sequence; In step S2, constructing the novel Token-Vocabulary serial single-modal intra-KAN feature interaction module for user-item interaction specifically includes the following steps: S21: Introduce the Token-level and Channel-level KAN neural network modules to enhance the expression ability and mutual understanding ability of modal information in the high-dimensional feature space; S22: Adopt an improved feature normalization mechanism to achieve feature consistency within a single modality and between cross-modalities through LayerNorm combined with the TokenKAN and ChannelKAN modules; For the visual modal feature sequence and the text modal feature sequence , the modal feature sequence of the identifier id , the function calculation formulas within each single modality are as follows: ; ; ; ; ; ; Among them, represents the visual modality information after the Token-level feature interaction of the visual modality features, represents the text modality information after the Token-level feature interaction of the text modality features, represents the identifier modality information after the Token-level feature interaction of the identifier modality features, i represents the information interaction in the column direction of the feature interaction matrix, represents the visual modality features after the Channel-level feature fusion of the visual modality information, represents the text modality features after the Channel-level feature fusion of the text modality information, represents the identifier modality features after the Channel-level feature fusion of the identifier modality information, j represents the information interaction in the row direction of the feature interaction matrix.
[0021] Step S3: Construct a Token-Channel serial multi-modal KAN feature interaction and fusion module for new user-item interaction, fully integrate multi-modal feature learning, perform multi-modal feature concatenation operations on each single-modal feature matrix, perform item feature interaction and fusion in the token dimension and the channel dimension respectively, model and learn cross-modal user preference information, and generate the final embedding learning representation; In step S3, constructing a Token-Channel serial multi-modal KAN feature interaction and fusion module for new user-item interaction specifically includes the following steps: S31: Introduce the feature outputs , and generated by the single-modal intra-feature interaction module based on the KAN network in step S2, perform multi-modal feature concatenation operations on each single-modal feature matrix, and obtain the multi-modal feature interaction matrix that needs to be fused. The function formula is as follows: ; Among them, represents the linear mapping dimensionality reduction module, which reduces the features of to ; represents a cascading splicing operation, which concatenates multi-source and multi-modal feature matrices; In step S31, for the complex feature interactions in multi-modal data, a feature fusion scheme based on dual optimization of Token dimension and Channel dimension is proposed to construct an item multi-modal feature interaction matrix, achieve efficient cooperation and information flow between multi-modalities, and dynamically adjust the feature interaction weights between modalities through the distributed weight learning of the KAN network.
[0022] S32: For the multi-modal feature matrix after cascading splicing, normalization is used in the marked dimension to limit the feature numerical size to prevent gradient disappearance and gradient explosion during the training process; In step S32, for the feature optimization strategy of differential alignment between modalities, effective feature fusion and denoising are achieved within and between modalities through a layer-by-layer interaction mechanism. The differences between modalities include vision, text, and identifier sequences.
[0023] S33: Use the KAN neural network fusion interaction operation between multi-modalities to fuse the feature information of the multi-modal information marked dimension, and use residual connection to inject the feature results before calculation to generate the intermediate step feature embedding representation results after fusion between multi-modalities , and the specific function formula is as follows: ; In the formula, represents layer normalization operation to prevent gradient disappearance and gradient explosion during the model training process, represents the multi-modal feature information after Token-level feature interaction, and i represents information interaction in the column direction of the feature interaction matrix; S34: Based on the intermediate step feature embedding representation results after multi-modal fusion , normalization operation is performed in the channel dimension, and the KAN neural network interaction fusion operation between multi-modalities in the channel dimension is executed to perform the interactive learning of features between modalities and generate the final feature embedding representation results , and the function formula is as follows: ; In the formula, represents the multi-modal feature information after Channel-level feature interaction, and j represents information interaction in the column direction of the feature interaction matrix.
[0024] In step S34, a progressive recommendation result generation mechanism is adopted, introducing L2 regularization and the layer-by-layer feature fusion strategy of the KAN network to improve the diversity and accuracy of the final recommendation results.
[0025] Step S4: For the embedded learning representation generated in step S3, use the item prediction layer prediction module to predict the next specific item to interact with the user. Obtain a lower-dimensional representation and classifier through a multi-layer perceptron to screen and subdivide the item, and obtain the recommendation result.
[0026] Embodiment As Figure 1 shown, a multi-modal sequence recommendation method based on the KAN network structure is provided to explicitly learn the information between multi-source and multi-modal heterogeneous data in the multi-modal sequence recommendation scenario, mine the fine-grained user interest preferences in different modalities of the user, fully integrate and learn the multi-modal interest representation, and capture and satisfy the user's potential intentions. It includes a multi-modal feature extraction module, a Token-Channel serial multi-modal KAN feature interaction and fusion module between the user and the item, a Token-Channel serial multi-modal KAN feature interaction and fusion module between the user and the item, and a final item identifier prediction module.
[0027] In the actual recommendation application scenario, in addition to the learnable identifier, it usually includes the text modality of the item (for example: description information, title, trademark, etc. of the product) and the image modality information (for example: legend of the product, etc.). Therefore, in this embodiment, the text modality and the image modality data are used as examples: Taking the image, text, and item identifier three-modal sequences in the user's historical interaction (such as: click, view, comment, etc.) record as the input, first use the corresponding modal feature extraction module to extract and process the sequence information of the image, text, and item for the three single-modal data respectively. Then execute the intra-modal feature fusion module to mine the user's fine-grained feature preferences for the single-modal information. It includes a normalization layer and a residual connection to prevent gradient disappearance and gradient explosion during the training process and enhance the stability of the training process. Subsequently, a multi-modal feature interaction and fusion module based on a post-processing fusion strategy is introduced for information interaction and fusion between multi-modalities, and comprehensively mine the potential interest preference information of the user in multi-modalities. For the finally fused user preference learning representation, finally input it into the final item identifier prediction module to predict the next item that the user may interact with, and then execute the personalized recommendation service.
[0028] In the multi-modal feature extraction module, take the given user historical interaction multi-modal item sequence as the input. First, for the modal sequence information such as vision and text, use a single-modal encoder and mapping module for feature extraction. For the identifier modal sequence information, use an identifier retriever to retrieve the learnable embedding representation of the corresponding item, and further construct the feature matrix under each modality. The embedding representations of each modality obtained by the multi-modal feature extraction module are , for use by subsequent modules.
[0029] In the Token-Channel serial multimodal KAN feature interaction and fusion module for user-item interaction, first, based on the KAN neural network, we introduce the single-modal internal feature interaction module built entirely based on the KAN structure. Compared with traditional multi-layer perceptrons, KAN does not use a fixed activation function on neuron nodes but uses a learnable activation function on the edges of the network structure during information transmission for non-linear activation transformation. This process uses an adaptive and learnable B-spline function for parameterization and optimization during training to avoid the computationally expensive linear weight multiplication operation in multi-layer perceptrons, thus achieving lightweight model parameter scale and a certain degree of interpretability. Compared with traditional multi-layer perceptrons based on the universal approximation theorem, KAN is a new neural network structure evolved from the Kolmogorov-Arnold representation theorem, which proves that any smooth multivariate continuous function , can be represented as a finite combination of univariate continuous functions and additive binary operations:
[0030] where , , . This shows that a bounded multivariate function can be represented as a sum of univariate functions, and learning a high-dimensional function can be reduced to learning the form of simple polynomial one-dimensional functions. However, in actual application scenarios, these one-dimensional functions may be non-smooth, so although reasonable in theory, they cannot be applied to actual application scenarios. KAN generalizes the application of the Kolmogorov-Arnold representation theorem, extends the two-layer non-linear summation in the above formula as the hidden layer representation to arbitrary widths and depths, and uses techniques such as backpropagation and pruning to be closer to actual applications than previous studies. At the same time, most functions are smooth and have a sparse compositional structure, so it can further contribute to smooth representation learning.
[0031] For a supervised learning task, the goal is to learn a function , to approximate the input-output sample pairs such that . Based on the above theoretical analysis, the solution content of this task can be decomposed into the process of seeking appropriate univariate functions and , thereby introducing a parameterized neural network to fit and Since both functions are single-variable functions, they can be learned by fitting each one-dimensional function parameterized as a B-spline curve. By analogy with the hierarchical structure of the traditional multi-layer perceptron (including linear transformation and non-linear activation functions), the hierarchical structure of KAN can be defined as follows: ; where, represents the dimension of the input features, represents the dimension of the output features, and the function contains trainable parameters. By stacking multiple layers of the KAN hierarchical structure, the shape of the final KAN network can be expressed as:
[0032] where, is the total number of nodes in the -th layer in the computational graph. We use to represent the -th neuron in the -th layer, and use to represent the activation value of this neuron. In KAN, there are a total of activation functions in the -th layer and the -th layer. The activation function connecting - neuron and -neuron can be expressed as:
[0033] where, represents the activation function of the activation value , and the result after activation is expressed as . Further, the activation value of the
[0034] -neuron is obtained by summing up all the results after activation of all neurons connected to this neuron:
[0035] where, the matrix on the right side of the equation is , which represents the function matrix corresponding to the -th layer of the KAN layer. Assuming a general KAN neural network structure consists of layers, when a given input vector is provided, the output result after passing through the KAN neural network can be expressed as:
[0036] Among them, represents an operation performed continuously. Correspondingly, a traditional multi-layer perceptron can be represented as a continuous linear transformation on the input vector and the non-linear activation function combination:
[0037] From this, it can be observed that compared with the traditional multi-layer perceptron, the KAN network uses a combination of linear transformation and non-linear activation function for substitution. In the KAN network, the specific function calculation formula is defined as:
[0038]
[0039] Among them, mainly contains basis functions to implement residual connection and spline function weighted sum of the two. Usually, parameterized as a linear combination of B-spline functions, where the weights of the linear combination are learnable:
[0040] Based on this, different from the traditional multi-layer perceptron, we design a single-modal intra-feature interaction module completely based on the KAN network structure for information interaction within the modal feature structure. Specifically, through the embedding representations of each modality obtained by each modality feature encoder and mapping module , first perform a normalization operation on the token dimension to accelerate model training and prevent gradient vanishing and gradient explosion, and then perform a single-modal KAN neural network interaction operation to perform intra-modal feature interaction learning, and use residual connection to optimize the network learning ability and solve the deep degradation problem. Generally speaking, the embedding representations of each modality generate the function formula of the intermediate step feature embedding representation result as follows: ; Specifically, for the visual modality, visual features are extracted through a visual encoder and a visual mapping module ( ), for text information, text features are extracted through a text encoder and a text mapping module ( ), for the identifier modality sequence information, a learnable one is retrieved using an identifier retriever Characterization ( ) Through the respective single-modal internal labeled dimension KAN neural network interaction operations, there are:
[0041]
[0042]
[0043] Based on the intermediate step feature embedding representation results of each modality , for the intermediate feature calculation results within each modality, perform a normalization operation on the channel dimension within each modality, and execute a single-modal internal KAN neural network interaction operation on the channel dimension to perform interactive learning of intra-modal features and generate intra-modal feature embedding representation results . The function formula is as follows: ; Specifically, for the intermediate feature calculation results of the visual modality ( ), the intermediate calculation results of the text modality ( ), and the intermediate calculation results of the learnable identifier modality ( ), the function formula is as follows: ; ; .
[0044] In the Token-Channel serial multi-modal KAN feature interaction fusion module for user-item interaction, for the , and of the feature outputs generated by each single-modal internal feature interaction module based on the KAN network, first perform a multi-modal feature concatenation operation on each single-modal feature matrix to obtain a multi-modal feature interaction matrix to be fused . The function formula is as follows: ; Among them, represents a linear mapping dimensionality reduction module used to reduce the features of to to reduce the computational complexity; Denotes a cascading splicing operation for concatenating multi-source multi-modal feature matrices. Through the concatenated multi-modal feature matrix, first, normalization is used in the token dimension to limit the feature value size to prevent gradient vanishing and gradient explosion during training. Second, the KAN neural network fusion interaction operation between multi-modalities is used to fuse the feature information in the token dimension of multi-modal information and use residual connections to inject the feature results before calculation to solve the problem of performance degradation of the model as the network depth increases. At the same time, the network learning ability is optimized, and finally, the intermediate step feature embedding representation result after multi-modal fusion is generated The function formula is as follows: ; Finally, based on the intermediate step feature embedding representation result after multi-modal fusion , normalization operation is performed in the channel dimension, and the KAN neural network interaction fusion operation between multi-modalities in the channel dimension is executed to perform interactive learning of features between modalities and generate the final feature embedding representation result , and the function formula is as follows: .
[0045] In the final item identifier prediction module, after passing through the layer of single-modal intra-feature interaction module based on the KAN network and the layer of multi-modal inter-feature interaction fusion module based on the KAN network structure, after obtaining the final feature embedding representation result, by combining it with the previous item embeddings of user interactions and all item candidate embeddings perform a similarity calculation operation. Finally, each item is screened and segmented by a classifier to obtain the final recommendation result. The function formula is as follows:
[0046] where maps each value in the classification result to the range of 0 and 1, and the elements are combined into 1 to generate a probability distribution representation of each recommended item. By using a multi-layer perceptron to obtain a lower-dimensional representation and a classifier to screen and segment item candidates, the final recommendation result is obtained.
[0047] Therefore, the present invention adopts the above-mentioned multi-modal sequence recommendation method based on the KAN network structure. By using the KAN network structure design, the computational complexity is greatly reduced, lightweight and high efficiency are achieved. By explicitly capturing the information relationship between multi-source multi-modal heterogeneous data, the fine-grained interest preferences of users in different modalities are deeply mined, and multi-modal interest representations are fully fused to better meet the potential intentions of users.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements do not enable the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multimodal sequence recommendation method based on KAN network structure, characterized by: The specific steps include: Step S1: Given a user's historical interaction multimodal item sequence as input, extract the corresponding modal features of the item sequence; Step S2: Construct a new Token-Channel serial unimodal intra-KAN feature interaction module for user-item interaction, optimize the modeling of users' fine-grained interest preferences under a single specific modality and interactive learning of the unimodal intra-representation of item sequences; Step S3: Construct a new Token-Channel serial multimodal KAN feature interaction fusion module for user-item interaction, fully integrate multimodal feature learning, perform multimodal feature cascade operations on each single-modal feature matrix, perform item feature interaction fusion in the tag dimension and channel dimension respectively, model and learn cross-modal user preference information, and generate the final embedded learning representation; Step S4: For the embedded learning representation generated in step S3, use the item prediction layer prediction module to predict the next specific item that the user will interact with. Use a multi-layer perceptron to obtain a lower-dimensional representation and classifier to screen and segment the items to obtain the recommendation results.
2. According to the multimodal sequence recommendation method based on the KAN network structure in claim 1, it is characterized by: In step S2, a new Token-Channel serial single-modal intra-KAN feature interaction module for user-item interaction is constructed, which specifically includes the following steps: S21: The introduction of Token-level and Channel-level KAN neural network modules improves the expression and mutual understanding capabilities of modal information in high-dimensional feature space; S22: An improved feature normalization mechanism is used to achieve feature consistency within a single modality and across modalities through LayerNorm combined with TokenKAN and ChannelKAN modules; For visual modality feature sequences , text modality feature sequence , the modal feature sequence of identifier id , the calculation formula of each single mode internal function is as follows: ; ; ; ; ; ; in, Represents the visual modal information after the visual modal features interact with the Token-level features. Represents the text modal information after the text modal features have been interacted with the Token-level features. It represents the identifier modal information after the identifier modal feature has been interacted with the Token-level features. i represents the information interaction in the column direction of the feature interaction matrix. It represents the visual modality features after the visual modality information is fused with the Channel-level features. It represents the text modality features after the text modality information is fused with the Channel-level features. It represents the identifier modal features after the identifier modal information is fused with the Channel-level features, and j represents the information interaction in the row direction of the feature interaction matrix.
3. According to the multimodal sequence recommendation method based on the KAN network structure in claim 1, it is characterized by: In step S3, a new Token-Channel serial multimodal KAN feature interaction fusion module for user-item interaction is constructed, which specifically includes the following steps: S31: Introduce the feature output generated by the single-modal intra-feature interaction module based on the KAN network in step S2 , and , perform multimodal feature cascade operations on each single-modal feature matrix to obtain the multimodal feature interaction matrix that needs to be fused , the function formula is as follows: ; in, Represents a linear mapping dimensionality reduction module, The feature dimension reduction is ; Represents the cascade concatenation operation, which connects multi-source and multi-modal feature matrices in series; S32: The multimodal feature matrix after cascading and splicing. Normalization is used in the label dimension to limit the feature value size to prevent gradient vanishing and gradient exploding during training. S33: Fusion of interactive operations using multimodal KAN neural networks The feature information of the dimension is fused with multimodal information, and the feature result before calculation is injected using residual connection to generate the intermediate step feature embedding representation result after multimodal fusion. , the specific function formula is as follows: ; In the formula, The normalization operation of the representation layer is used to prevent the gradient from disappearing and exploding during the model training process. It represents the multimodal feature information after the token-level feature interaction, and i represents the information interaction in the column direction of the feature interaction matrix; S34: Intermediate step feature embedding representation results based on multimodal fusion , using normalization operation in the channel dimension, performing multi-modal KAN neural network interactive fusion operation in the channel dimension , to interactively learn features between modalities and generate the final feature embedding representation result , the function formula is as follows: ; In the formula, It represents the multimodal feature information after the Channel-level feature interaction, and j represents the information interaction in the column direction of the feature interaction matrix.
4. According to the multimodal sequence recommendation method based on the KAN network structure described in claim 3, it is characterized by: In step S31, in response to the complex feature interactions in multimodal data, a feature fusion scheme based on dual optimization of Token dimension and Channel dimension is proposed to construct a multimodal feature interaction matrix of items, realize efficient collaboration and information flow among multimodalities, and dynamically adjust the feature interaction weights among modalities through distributed weight learning of KAN network.
5. According to the multimodal sequence recommendation method based on the KAN network structure as described in claim 3, it is characterized by: In step S32, a feature optimization strategy for inter-modality difference alignment is used to achieve effective feature fusion and denoising within and between modalities through a layer-by-layer interaction mechanism. The inter-modality differences include vision, text, and identifier sequences.
6. According to the multimodal sequence recommendation method based on the KAN network structure as described in claim 3, it is characterized by: In step S34, a progressive recommendation result generation mechanism is adopted, and L2 regularization and a layer-by-layer feature fusion strategy of the KAN network are introduced to improve the diversity and accuracy of the final recommendation results.
Citation Information
Patent Citations
Graph Transform article recommendation method based on multi-modal semantic fusion
CN118657560A
Group video recommendation system and method fusing multi-modal information and interest similarity
CN118690037A
Graph neural network multi-modal recommendation method and system based on feature redundancy removal
CN118760804A
Scene adaptive video compression method and system based on natural language guidance
CN118972590A
Self-adaptive alignment cross-modal vision-language ship intelligent man-machine interaction method
CN119357897A