A Multimodal Sequential Recommendation Method Based on the KAN Network Structure
Through the multimodal sequence recommendation method based on the KAN network structure, the fine-grained interest preferences of modeling users under different modes are optimized, and the problems of high computational complexity and low efficiency in the prior art are solved, and efficient personalized recommendation is achieved.
Patent Information
- Application Number
- CN202510534576.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing multimodal sequence recommendation system framework has high computational complexity and low space utilization efficiency, making it difficult to effectively explore users' preference information under different modes.
Using a multimodal sequence recommendation method based on KAN network structure, the Token-Channel serial single-modal and multimodal KAN feature interaction module is constructed to optimize the fine-grained interest preferences of modeling users under different modes, and a lightweight design is adopted to reduce computing complexity and improve system operation efficiency.
It significantly reduces the computational complexity, improves the system operation efficiency, deeply explores users' fine-grained interests and preferences under multimodality, and achieves more accurate and personalized recommendations.
Smart Images

Figure CN120067460B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and particularly to a multi-modal sequence recommendation method based on a KAN network structure. Background Art
[0002] With the booming development of Internet multimedia sharing platforms, recommendation systems have become an indispensable core component. A recommendation system aims to mine potential preference intention information of users based on users' historical behavior records, and provide personalized recommendation services for different users in an implicit manner. Compared with traditional recommendation algorithms based on the id paradigm, the research on multi-modal recommendation algorithms is still in its infancy. Multi-modal recommendation algorithms mine potential user preference information of users in different modalities by means of the item sequences and multi-modal attribute features of items with which users have interacted historically.
[0003] However, existing multi-modal sequence recommendation system frameworks generally have problems such as high time computational complexity, low space utilization efficiency, and difficulty in effectively mining user preference information in each modality. Summary of the Invention
[0004] To solve the above problems, the present invention provides a multi-modal sequence recommendation method based on a KAN network structure, which explicitly models the information association between multi-source multi-modal heterogeneous data, deeply mines the fine-grained interest preferences of users in different modalities, fully integrates multi-modal interest representations, and more accurately meets the potential needs of users. At the same time, the framework adopts a lightweight KAN network structure design, significantly reducing the computational complexity and improving the system operation efficiency, and having wide practical application value.
[0005] To achieve the above object, the present invention provides a multi-modal sequence recommendation method based on a KAN network structure, specifically including the following steps:
[0006] Step S1: Given a user's historical interactive multi-modal item sequence as input, extract the corresponding modal features of the item sequence;
[0007] Step S2: Construct a Token-Channel serial single-modal intra-KAN feature interaction module for new user-item interaction, and optimize the modeling of the user's fine-grained interest preferences under a single specific modality and the interaction learning of the single-modal intra-representation of the item sequence;
[0008] Step S3: Construct a Token-Channel serial multi-modal inter-KAN feature interaction and fusion module for new user-item interaction, fully integrate multi-modal feature learning, perform multi-modal feature concatenation operations on each single-modal feature matrix, perform item feature interaction and fusion respectively in the token dimension and the channel dimension, model and learn cross-modal user preference information, and generate a final embedding learning representation;
[0009] Step S4: For the embedded learning representation generated in step S3, use the item prediction layer prediction module to predict the next specific item to interact with the user. Obtain a lower-dimensional representation and classifier through a multi-layer perceptron to screen and subdivide the item, and obtain the recommendation result.
[0010] Preferably, in step S2, constructing a new user-item interaction labeled-word table serial single-modal intra-KAN feature interaction module specifically includes the following steps:
[0011] S21: Introduce the KAN neural network module at the Token level and the Channel level, which improves the expression ability and mutual understanding ability of modal information in the high-dimensional feature space;
[0012] S22: Adopt an improved feature normalization mechanism to achieve feature consistency within and across modalities through LayerNorm combined with the TokenKAN and ChannelKAN modules;
[0013] For the visual modality feature sequence and the text modality feature sequence and the modality feature sequence of the identifier id , the calculation formulas of each single-modal function are as follows:
[0014] ;
[0015] ;
[0016] ;
[0017] ;
[0018] ;
[0019] ;
[0020] Among them, represents the visual modality information after Token-level feature interaction of the visual modality feature, represents the text modality information after Token-level feature interaction of the text modality feature, represents the identifier modality information after Token-level feature interaction of the identifier modality feature, and i represents information interaction in the column direction of the feature interaction matrix. represents the visual modality feature after Channel-level feature fusion of the visual modality information, represents the text modality feature after Channel-level feature fusion of the text modality information, It represents the identifier modality feature after the identifier modality information undergoes Channel-level feature fusion, and j represents the information interaction in the row direction of the feature interaction matrix.
[0021] Preferably, in step S3, a Token-Channel serial multi-modal KAN feature interaction fusion module for new user-item interaction is constructed, which specifically includes the following steps:
[0022] S31: Introduce the feature output generated by the single-modal intra-feature interaction module based on the KAN network in step S2 、 and , perform a multi-modal feature concatenation operation on each single-modal feature matrix to obtain a multi-modal feature interaction matrix that needs to be fused , and the function formula is as follows:
[0023] ;
[0024] Among them, represents a linear mapping dimensionality reduction module that reduces the features of to ; represents a concatenation operation that concatenates multi-source multi-modal feature matrices;
[0025] S32: For the concatenated multi-modal feature matrix, use normalization in the marked dimension to limit the feature value size to prevent gradient vanishing and gradient explosion during the training process;
[0026] S33: Use the multi-modal KAN neural network fusion interaction operation to fuse the feature information in the marked dimension of multi-modal information, and use residual connection to inject the feature result before calculation to generate the intermediate step feature embedding representation result after multi-modal fusion , and the specific function formula is as follows:
[0027] ;
[0028] In the formula, represents the layer normalization operation to prevent gradient vanishing and gradient explosion during the model training process, represents the multi-modal feature information after Token-level feature interaction, and i represents the information interaction in the column direction of the feature interaction matrix.
[0029] S34: Based on the intermediate step feature embedding representation result after multi-modal fusion , perform a normalization operation in the channel dimension and execute the multi-modal KAN neural network interaction fusion operation in the channel dimension , for interactive learning of features between modalities to generate the final feature embedding representation result , and the function formula is as follows:
[0030] ;
[0031] In the formula, represents the multi-modal feature information after Channel-level feature interaction, and j represents information interaction in the column direction of the feature interaction matrix.
[0032] Preferably, in step S31, for the complex feature interaction in multi-modal data, a feature fusion scheme based on dual optimization of Token dimension and Channel dimension is proposed to construct an item multi-modal feature interaction matrix, realizing efficient cooperation and information flow between multi-modalities. Through the distributed weight learning of the KAN network, the feature interaction weights between modalities are dynamically adjusted.
[0033] Preferably, in step S32, for the feature optimization strategy of inter-modal difference alignment, effective feature fusion and denoising are realized within and between modalities through a layer-by-layer interaction mechanism. The differences between modalities include vision, text, and identifier sequences.
[0034] Preferably, in step S34, a progressive recommendation result generation mechanism is adopted, introducing L2 regularization and the layer-by-layer feature fusion strategy of the KAN network to improve the diversity and accuracy of the final recommendation result.
[0035] Therefore, the present invention adopts the above multi-modal sequence recommendation method based on the KAN network structure, having the following beneficial effects:
[0036] (1) In the present invention, the KAN neural network is introduced as the basic computing unit, optimizing the neural network structure of the traditional attention calculation mechanism with quadratic time complexity, enhancing the discriminative ability of the representation, and improving the efficiency of the model.
[0037] (2) The present invention designs a new Token-Channel serial multi-modal KAN feature interaction fusion module for user-item interaction, intending to interact and fuse the information within specific single modalities from two perspectives of the token dimension and the channel dimension for each modality of the user-interacted item sequence, and mine the fine-grained potential preference intention information of the user in each modality.
[0038] (3) The present invention designs a new Token-Channel serial multi-modal KAN feature interaction fusion module for user-item interaction, and then performs feature interaction and fusion in the token dimension and the channel dimension, effectively modeling cross-modal user preference information, generating the final embedding representation, and deeply mining the behavior preference characteristics of the user under multi-modal conditions.
[0039] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0040] Figure 1 It is a schematic diagram of a multi-modal sequence recommendation method based on the KAN network structure of the present invention;
[0041] Figure 2 It is a specific detailed structure diagram inside the feature interaction and fusion module based on the KAN network structure in the present invention. Detailed Embodiment
[0042] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0043] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the field to which the present invention belongs.
[0044] The words such as "including" or "comprising" used in the present invention mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements. The orientation or positional relationship indicated by terms such as "inside", "outside", "above", "below", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation to the present invention. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly. In the present invention, unless otherwise clearly defined and limited, terms such as "attached" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0045] A multi-modal sequence recommendation method based on the KAN network structure specifically includes the following steps:
[0046] Step S1: Given the user's historical interaction multi-modal item sequence as input, extract the corresponding modal features of the item sequence;
[0047] Step S2: Construct a Token-Channel serial single-modal intra-KAN feature interaction module for the interaction between the new user and the item, and optimize the modeling of the user's fine-grained interest preference under a single specific modality and the interaction learning of the single-modal intra-representation of the item sequence;
[0048] In step S2, constructing the labeled-word list serial single-modal inner KAN feature interaction module for the interaction between new users and items specifically includes the following steps:
[0049] S21: Introduce the KAN neural network modules at the Token level and the Channel level, which enhances the expression ability and mutual understanding ability of modal information in the high-dimensional feature space;
[0050] S22: Adopt an improved feature normalization mechanism to achieve feature consistency within a single modality and across modalities through LayerNorm combined with the TokenKAN and ChannelKAN modules;
[0051] For the visual modality feature sequence and the text modality feature sequence and the modality feature sequence of the identifier id , the function calculation formulas within each single modality are as follows:
[0052] ;
[0053] ;
[0054] ;
[0055] ;
[0056] ;
[0057] ;
[0058] Among them, represents the visual modality information after Token-level feature interaction of the visual modality feature, represents the text modality information after Token-level feature interaction of the text modality feature, represents the identifier modality information after Token-level feature interaction of the identifier modality feature, i represents information interaction in the column direction of the feature interaction matrix, represents the visual modality feature after Channel-level feature fusion of the visual modality information, represents the text modality feature after Channel-level feature fusion of the text modality information, represents the identifier modality feature after Channel-level feature fusion of the identifier modality information, j represents information interaction in the row direction of the feature interaction matrix.
[0059] Step S3: Construct a Token-Channel serial multi-modal KAN feature interaction and fusion module for new user-item interaction, fully integrate multi-modal feature learning, perform multi-modal feature concatenation operations on each single-modal feature matrix, perform item feature interaction and fusion respectively in the token dimension and the channel dimension, model and learn cross-modal user preference information, and generate the final embedded learning representation;
[0060] In step S3, construct a Token-Channel serial multi-modal KAN feature interaction and fusion module for new user-item interaction, which specifically includes the following steps:
[0061] S31: Introduce the feature output generated by the single-modal intra-feature interaction module based on the KAN network in step S2 、 And , perform multi-modal feature concatenation operations on each single-modal feature matrix to obtain a multi-modal feature interaction matrix that needs to be fused , and the function formula is as follows:
[0062] ;
[0063] Among them, represents a linear mapping dimensionality reduction module that reduces the features of to ; represents a concatenation operation that concatenates multi-source multi-modal feature matrices;
[0064] In step S31, for the complex feature interactions in multi-modal data, a feature fusion scheme based on dual optimization of the token dimension and the channel dimension is proposed, an item multi-modal feature interaction matrix is constructed, efficient cooperation and information flow between multi-modalities are achieved, and the feature interaction weights between modalities are dynamically adjusted through the distributed weight learning of the KAN network.
[0065] S32: For the multi-modal feature matrix after concatenation, use normalization in the token dimension to limit the feature numerical size to prevent gradient disappearance and gradient explosion during training;
[0066] In step S32, for the feature optimization strategy of cross-modal difference alignment, effective feature fusion and denoising are achieved within and between modalities through a layer-by-layer interaction mechanism. The differences between modalities include vision, text, and identifier sequences.
[0067] S33: Use the KAN neural network for cross-modal fusion interaction operation to fuse the feature information in the multi-modal information token dimension, and use residual connection to inject the feature results before calculation to generate the intermediate step feature embedding representation result after cross-modal fusion , the specific function formula is as follows:
[0068] ;
[0069] In the formula, represents layer normalization to prevent gradient vanishing and gradient explosion during model training, represents the multi-modal feature information after Token-level feature interaction. i represents information interaction in the column direction of the feature interaction matrix;
[0070] S34: Based on the intermediate step feature embedding representation result after multi-modal fusion , perform a normalization operation in the channel dimension and execute the cross-modal KAN neural network interaction fusion operation in the channel dimension , to perform cross-modal feature interaction learning and generate the final feature embedding representation result , and the function formula is as follows:
[0071] ;
[0072] In the formula, represents the multi-modal feature information after Channel-level feature interaction. j represents information interaction in the column direction of the feature interaction matrix.
[0073] In step S34, a progressive recommendation result generation mechanism is adopted, introducing L2 regularization and the layer-by-layer feature fusion strategy of the KAN network to improve the diversity and accuracy of the final recommendation result.
[0074] Step S4: For the embedding learning representation generated in step S3, use the item prediction layer prediction module to predict the specific item to interact with the user next. Obtain a lower-dimensional representation through a multi-layer perceptron and use a classifier to screen and segment the item to obtain the recommendation result.
[0075] Embodiment
[0076] As Figure 1 shown, a multi-modal sequence recommendation method based on the KAN network structure is provided to explicitly learn the information between multi-source multi-modal heterogeneous data in the multi-modal sequence recommendation scenario, mine the fine-grained user interest preferences of users in different modalities, fully fuse and learn multi-modal interest representations, and capture and satisfy the potential intentions of users. It includes a multi-modal feature extraction module, a Token-Channel serial cross-modal KAN feature interaction fusion module for user-item interaction, a Token-Channel serial cross-modal KAN feature interaction fusion module for user-item interaction, and a final item identifier prediction module.
[0077] In actual recommendation application scenarios, in addition to learnable identifiers, the items usually include text modality (e.g., description information, title, trademark, etc. of the product) and image modality information (e.g., legend of the product, etc.). Therefore, in this embodiment, text modality and image modality data are used as examples:
[0078] According to the three modal sequences of images, texts and item identifiers in the user's historical interaction records (such as clicks, views, comments, etc.) as input, the corresponding modal feature extraction modules are first used to extract and process the sequence information of images, texts and items for the three unimodal data. Then, the intra-unimodal feature fusion module is executed to mine the user's fine-grained feature preferences for unimodal information. It contains normalization layers and residual connections to prevent gradient vanishing and gradient explosion during training and enhance the stability of the training process. Subsequently, a multimodal feature interaction fusion module based on a post-processing fusion strategy is introduced to interact and fusion information between multiple modalities, and comprehensively mine the user's potential interest preference information under multiple modalities. The final fused user preference learning representation is finally input into the final item identifier prediction module to predict the next item that the user may interact with, and then perform personalized recommendation services.
[0079] In the multimodal feature extraction module, the given user historical interaction multimodal item sequence is taken as input. First, the single modality encoder and mapping module are used to extract features for the visual, textual and other modal sequence information. The identifier retriever is used to retrieve the learnable embedding representation of the corresponding item for the identifier modal sequence information. The feature matrix under each modality is further constructed. The embedding representation of each modality obtained by the multimodal feature extraction module is: , for use by subsequent modules.
[0080] In the Token-Channel serial multimodal KAN feature interaction fusion module for user-item interaction, we first derive the single-modal intra-feature interaction module that is completely based on the KAN structure based on the KAN neural network. Compared with traditional multi-layer perceptrons, KAN does not use fixed activation functions on neuron nodes, but uses learnable activation functions to perform nonlinear activation transformations on the edges of the network structure during information transmission. This process uses adaptive learnable B-spline functions for parameterization and optimization during training to avoid the linear weight multiplication operations with high learning costs in multi-layer perceptrons, thereby achieving lightweight model parameter scale and a certain degree of interpretability. Compared with traditional multi-layer perceptrons based on universal approximation theory, KAN is a new neural network structure evolved based on the Kolmogorov-Arnold representation theorem, which proves that smooth multivariate continuous functions on any bounded domain , can all be expressed as a finite combination of univariate continuous functions and additive binary operations:
[0081]
[0082] Among them , , . It shows that bounded multivariate functions can be expressed in the form of summation of univariate functions, and learning high-dimensional functions can be reduced to learning the form of simple polynomial univariate functions. However, in actual application scenarios, these univariate functions may be non-smooth, so it is reasonable in theory but cannot be applied to actual application scenarios. KAN generalizes the application of the Kolmogorov-Arnold representation theorem, generalizes the two-layer nonlinear summation in the above formula as the hidden layer representation to arbitrary widths and depths, and uses techniques such as backpropagation and pruning to make it closer to actual applications than previous studies. At the same time, most functions are smooth and have a sparse compositional structure, so it can further contribute to smooth representation learning.
[0083] For a supervised learning task, the goal is to learn a function , to approximate the function for the input-output sample pairs , such that . Based on the above theoretical analysis, the solution content of this task can be decomposed into the process of seeking appropriate univariate functions and , thereby introducing a parametric neural network to fit and . Since both functions are univariate functions, they can be fitted and learned by parameterizing each one-dimensional function as a B-spline curve. By analogy with the hierarchical structure of traditional multi-layer perceptrons (including linear transformations and non-linear activation functions), the hierarchical structure of KAN can be defined as follows:
[0084] ;
[0085] Among them, represents the dimension of the input features, represents the dimension of the output features, and the function contains trainable parameters. By stacking multiple layers of the KAN hierarchical structure, the shape of the final KAN network can be expressed as:
[0086]
[0087] Among them, is the total number of nodes in the -th layer in the computational graph. We use to represent the -th layer and the neurons, and use to represent the activation value of this neuron. In the KAN, there are layers and layers, with a total of activation functions. The connection - neurons and - The activation function of the neuron can be expressed as:
[0088]
[0089] where represents the activation function of the activation value , and the result after activation is expressed as . Further, - The activation value of the neuron is obtained by summing up all the results after activation of all the neurons connected to this neuron:
[0090]
[0091] Rearranging the above formula into matrix form gives:
[0092]
[0093] where the matrix on the right side of the equation is , which represents the function matrix corresponding to the layer of the KAN layer. Assume that a general KAN neural network structure consists of layers. When a given input vector is provided, the output result of the KAN neural network can be expressed as:
[0094]
[0095] where represents the consecutive operations. Correspondingly, the traditional multi-layer perceptron can be expressed as the combination of consecutive linear transformation of the input vector and the non-linear activation function :
[0096]
[0097] It can be observed from this that, compared with the traditional multi-layer perceptron, the KAN network uses the combination of linear transformation and non-linear activation functions for substitution. In the KAN network, the specific function calculation formula is defined as:
[0098]
[0099]
[0100] Among them, mainly includes basis functions used to implement residual connection and spline functions and the weighted sum of the two. Usually, it is parameterized as a linear combination of B-spline functions, where the weights of the linear combination are learnable:
[0101]
[0102] Based on this, different from traditional multi-layer perceptrons, we design a single-modal intra-feature interaction module completely based on the KAN network structure for information interaction within the modal feature structure. Specifically, through the embedding representations of each modality obtained by each modality feature encoder and mapping module , first, a normalization operation is used in the token dimension to accelerate model training and prevent gradient vanishing and gradient explosion. Second, a single-modal intra-KAN neural network interaction operation is performed to conduct intra-modal feature interaction learning, and residual connections are used to optimize the network learning ability and solve the deep degradation problem. Generally speaking, the embedding representations of each modality generate the intermediate-step feature embedding representation results and the function formula is as follows:
[0103] ;
[0104] Specifically, for the visual modality, visual features ( ) are extracted through the visual encoder and visual mapping module. For text information, text features ( ) are extracted through the text encoder and text mapping module. For the identifier modality sequence information, a learnable representation ( ) is retrieved using the identifier retriever. Through the respective single-modal intra-token dimension KAN neural network interaction operations, there are:
[0105]
[0106]
[0107]
[0108] Based on the intermediate-step feature embedding representation results of each modality , for the calculation results of the intermediate features within each modality, perform a normalization operation in the channel dimension within each modality, and execute the single-modality KAN neural network interaction operation in the channel dimension , to perform interactive learning of the features within the modality and generate the feature embedding representation results within each modality , the function formula is as follows:
[0109] ;
[0110] Specifically, for the calculation results of the intermediate features of the visual modality ( ), the calculation results of the intermediate features of the text modality ( ), and the calculation results of the intermediate features of the learnable identifier modality ( ), the function formula is as follows:
[0111] ;
[0112] ;
[0113] .
[0114] In the Token-Channel serial multi-modal KAN feature interaction and fusion module for user-item interaction, for the , and of the feature outputs generated by each single-modal feature interaction module based on the KAN network, first perform a multi-modal feature concatenation operation on each single-modal feature matrix to obtain the multi-modal feature interaction matrix to be fused , the function formula is as follows:
[0115] ;
[0116] Among them, represents a linear mapping dimensionality reduction module, used to reduce the features of to to reduce the computational complexity; represents a concatenation operation, used to concatenate multi-source multi-modal feature matrices. Through the concatenated multi-modal feature matrix, first use normalization in the token dimension to limit the feature value size to prevent gradient disappearance and gradient explosion during the training process, and then use the multi-modal KAN neural network fusion interaction operation to fuse the feature information in the multi-modal information token dimension, and use residual connection to inject the feature results before calculation to solve the performance degradation problem of the model as the network depth increases, and at the same time optimize the network learning ability, and finally generate the intermediate step feature embedding representation results after multi-modal fusion The function formula is as follows:
[0117] ;
[0118] Finally, based on the intermediate step feature embedding representation result after multimodal fusion , a normalization operation is performed in the channel dimension, and the cross-modal KAN neural network interaction fusion operation in the channel dimension is executed , to perform cross-modal feature interaction learning and generate the final feature embedding representation result , the function formula is as follows:
[0119] .
[0120] In the final item identifier prediction module, after passing through layers of the single-modal intra-feature interaction module based on the KAN network and layers of the cross-modal feature interaction fusion module based on the KAN network structure, the final feature embedding representation result is obtained. Then, by comparing it with the item embeddings of the previous user interactions and all item candidate embeddings , a similarity calculation operation is performed. Finally, each item is screened and segmented by a classifier to obtain the final recommendation result. The function formula is as follows:
[0121]
[0122] where each value in the classification result is mapped to the range of 0 and 1, and all elements are combined to 1 to generate the probability distribution representation of each recommended item. By using a multi-layer perceptron to obtain a lower-dimensional representation and a classifier to screen and segment item candidates, the final recommendation result is obtained.
[0123] Therefore, the present invention adopts the above-mentioned multi-modal sequence recommendation method based on the KAN network structure. By using the KAN network structure design, the computational complexity is significantly reduced, achieving lightweight and high efficiency. By explicitly capturing the information relationships between multi-source and multi-modal heterogeneous data, the fine-grained interest preferences of users in different modalities are deeply mined, and multi-modal interest representations are fully fused to better meet the potential intentions of users.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements do not make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-modal sequence recommendation method based on the KAN network structure, characterized in that: Specifically, it includes the following steps: Step S1: Given the user's historical interaction multi-modal item sequence as input, extract the corresponding modal features of the item sequence; Step S2: Construct a novel Token-Channel serial single-modal intra-KAN feature interaction module for user-item interaction, optimize the modeling of the user's fine-grained interest preference under a single specific modality, and perform interactive learning on the single-modal representation of the item sequence; In step S2, constructing a novel Token-Channel serial single-modal intra-KAN feature interaction module for user-item interaction specifically includes the following steps: S21: Introduce Token-level and Channel-level KAN neural network modules to enhance the expression ability and mutual understanding ability of modal information in the high-dimensional feature space; S22: Adopt an improved feature normalization mechanism to achieve feature consistency within and across modalities through LayerNorm combined with TokenKAN and ChannelKAN modules; For the visual modality feature sequence and the text modality feature sequence and the modality feature sequence of the identifier id The calculation formula of the function within each single modality is as follows: ; ; ; ; ; ; Among them, represents the visual modality information after the Token-level feature interaction of the visual modality features, represents the text modality information after the Token-level feature interaction of the text modality features, represents the identifier modality information after the Token-level feature interaction of the identifier modality features, where i represents the information interaction in the column direction of the feature interaction matrix, represents the visual modality features after the Channel-level feature fusion of the visual modality information, represents the text modality features after the Channel-level feature fusion of the text modality information, represents the identifier modality features after the Channel-level feature fusion of the identifier modality information, where j represents the information interaction in the row direction of the feature interaction matrix; Step S3: Construct a novel Token-Channel serial multi-modal inter-KAN feature interaction and fusion module for user-item interaction, fully integrate multi-modal feature learning, perform multi-modal feature concatenation operations on each single-modal feature matrix, perform item feature interaction and fusion respectively in the token dimension and the channel dimension, model and learn cross-modal user preference information, and generate the final embedding learning representation; Step S4: For the embedding learning representation generated in step S3, use the item prediction layer prediction module to predict the next specific item to interact with the user, obtain a lower-dimensional representation through a multi-layer perceptron, and use a classifier to screen and segment the item to obtain the recommendation result.
2. The multimodal sequence recommendation method based on the KAN network structure according to claim 1, characterized in that: In step S3, constructing a novel Token-Channel serial multi-modal inter-KAN feature interaction and fusion module for user-item interaction specifically includes the following steps: S31: Introduce the feature output generated by the single-modal intra-feature interaction module based on the KAN network in step S2 , With , perform a multi-modal feature concatenation operation on each single-modal feature matrix to obtain a multi-modal feature interaction matrix that needs to be fused , and the function formula is as follows: , Among them, represents a linear mapping dimensionality reduction module that reduces the features to ; represents a cascading splicing operation that concatenates multi-source multi-modal feature matrices; S32: Concatenate the multi-modal feature matrices after concatenation, and use normalization in the token dimension to limit the feature numerical size to prevent gradient vanishing and gradient explosion during the training process; S33: Use the KAN neural network for multimodal fusion to perform interactive operations Fuse the feature information of the multimodal information marking dimension, and use residual connection to inject the feature results before calculation to generate the intermediate step feature embedding representation result after multimodal fusion , and the specific function formula is as follows: ; Wherein, represents layer normalization operation to prevent gradient vanishing and gradient explosion during the model training process, represents multi-modal feature information after Token-level feature interaction, and i represents information interaction in the column direction of the feature interaction matrix; S34: Intermediate step feature embedding representation result based on multi-modal fusion , perform a normalization operation in the channel dimension and execute the cross-modal KAN neural network interaction fusion operation in the channel dimension , to conduct cross-modal feature interaction learning and generate the final feature embedding representation result , and the function formula is as follows: ; In the formula, represents the multi-modal feature information after channel-level feature interaction, and j represents information interaction in the column direction of the feature interaction matrix.
3. A multi-modal sequence recommendation method based on the KAN network structure according to claim 2, characterized in that: In step S31, for the complex feature interaction in multi-modal data, a feature fusion scheme based on dual optimization of the token dimension and the channel dimension is proposed, an item multi-modal feature interaction matrix is constructed, efficient cooperation and information flow between modalities are realized, and the feature interaction weights between modalities are dynamically adjusted through the distributed weight learning of the KAN network.
4. A multi-modal sequence recommendation method based on the KAN network structure according to claim 2, characterized in that: In step S32, for the feature optimization strategy of cross-modal difference alignment, effective feature fusion and denoising are achieved within and across modalities through a layer-by-layer interaction mechanism, and the cross-modal differences include vision, text, and identifier sequences.
5. The multimodal sequence recommendation method based on the KAN network structure according to claim 2, wherein: In step S34, a progressive recommendation result generation mechanism is adopted, and L2 regularization and the layer-by-layer feature fusion strategy of the KAN network are introduced to improve the diversity and accuracy of the final recommendation result.
Citation Information
Patent Citations
Group video recommendation system and method fusing multi-modal information and interest similarity
CN118690037A
Sequence recommendation method based on multi-modal feature fusion
CN119783044A