A personalized recommendation method and system based on a large model

By using multi-channel hardware separation and large model analysis, independent feature packages are generated and a cross-modal similarity table is constructed. Combined with historical behavior records, dynamic compensation coefficients are generated to adjust node connection strength, solving the problems of cross-modal data bias and user interest drift in existing technologies, and realizing high-precision and real-time personalized recommendations.

CN120804432BActive Publication Date: 2025-11-18LUSTER LIGHTWAVE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511309604.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-11-18
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Existing large-model-based recommendation schemes have shortcomings in dynamically compensating for cross-modal data bias, capturing user interest drift, and deeply fusing multimodal information. This results in recommendation results deviating from the true distribution, failing to respond to user changes in a timely manner, and lacking deep fusion of non-textual modal features.

Method used

Multimodal content data streams are decomposed using multi-channel hardware separation to generate independent feature packages. A similarity correspondence table of cross-modal semantic relationships is established using a large model. By combining historical behavior records to analyze the sources of deviation, dynamic compensation coefficients are generated to adjust the connection strength between content nodes and user nodes, constructing an association expression vector. Finally, decision vectors are generated by integrating user purpose, content association, and timeliness features.

Benefits of technology

It achieves high-precision matching and ranking of cross-modal content and users' dynamic interests, improving the accuracy and timeliness of personalized recommendations and significantly alleviating the problems of recommendation lag and modal fragmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804432B_ABST
    Figure CN120804432B_ABST
Patent Text Reader

Abstract

The application provides a personalized recommendation method and system based on a large model. In the application, the content data stream is disassembled into independent feature packages of different modalities by a multi-channel acquisition server. The cross-modal semantic relationship between each feature package is analyzed by using a large model to establish a cross-modal similarity correspondence table. Combined with the cross-modal conversion records in the historical behavior, a dynamic compensation coefficient is generated through bias analysis. A bipartite graph structure of content and user is constructed, the compensation coefficient is converted into a connection strength adjustment factor, and the adjacent node weight generation table is integrated to generate an expression vector. Finally, the user purpose, content association and time value characteristic values of the vector are extracted, a dynamic weight model is constructed by combining the preference change trajectory, multi-feature integration is generated to generate a decision vector, cross-modal matching sorting is realized, and personalized recommendation results are output. The application realizes high-precision matching recommendation of multi-modal content and user demand by dynamically compensating the cross-modal data bias and fusing the user trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cross-modal large model recommendation methods, in particular to a personalized recommendation method and system based on a large model. BACKGROUND

[0002] In the multi-source heterogeneous content platform scenario, a large number of users continuously generate interactive content containing multiple modalities such as text, images, videos, and audio. These content sources are diverse and have significant structural differences, which poses a core demand for personalized recommendation systems: it must efficiently integrate multi-source heterogeneous data to build a unified representation, accurately understand the deep semantic associations between different modal content to bridge the cross-modal gap, and capture the dynamic changes in user interest preferences to address the cold start and interest drift problems, and ultimately achieve high-precision matching of content and user needs.

[0003] A current representative solution is a recommendation framework based on data augmentation of large language models combined with graph neural networks. The core process is to use a large language model to analyze user historical behavior and item text information, generate simulated user interaction records and fine-grained portrait labels to expand sparse data, and supplement missing attributes of items. Then, through a graph neural network, the relationship between users and items is modeled, and techniques such as noise pruning are used to improve the reliability of the generated data, and finally the recommendation results are output. This solution has shown certain effectiveness in improving coverage using text information.

[0004] Although the above-mentioned solution enhances the use of text data, it still has obvious limitations. The generated data may deviate from the platform's true distribution and introduce implicit noise, rely on offline batch processing mechanisms, and cause recommendation lag due to the inability to respond to the latest user dynamics in real time. The high cost of large model inference limits the scalability of large-scale item libraries, and lacks deep integration capabilities for non-text modalities such as image style or video dynamic features. Cross-modal alignment often relies on artificial rules, making it difficult to fully exploit the collaborative value of multi-source heterogeneous information. SUMMARY

[0005] The present application provides a personalized recommendation method and system based on a large model to solve the problem of insufficient dynamic compensation for cross-modal data bias, capturing user interest drift, and deep integration of multi-modal information in the prior art based on large model enhancement.

[0006] In a first aspect, the present application provides a personalized recommendation method based on a large model, comprising:

[0007] Obtaining a content data stream including a mixture of multi-modal features, configuring a multi-channel acquisition server, using the hardware separation unit of the multi-channel acquisition server to disassemble the content data stream into original data streams of different modal features, and generating independent feature packages according to the corresponding feature data and labels of the original data streams.

[0008] establish a cross-modal similarity correspondence table by a pre-established large model, analyzing the semantic relationship between the text features and the picture features among multiple independent feature packages, and the matching degree between the sound features and the video content;

[0009] determine the historical behavior records between different modal contents in the content data stream, and generate a dynamic compensation coefficient by bias source analysis combining the cross-modal conversion records in the historical behavior records and the similarity correspondence table;

[0010] Construct a bipartite graph structure of content nodes and user nodes, convert the dynamic compensation coefficient into an adjustment factor of connection strength, adjust the connection strength of the content nodes and user nodes, and integrate the correlation weight of adjacent nodes to generate an expression vector containing correlation features;

[0011] Extract the user purpose feature value, content contact feature value and time efficiency feature value in the expression vector, construct a dynamic change weight model combining the preference change trajectory of the historical behavior records, integrate the user purpose feature value, content contact feature value, time efficiency feature value and change weight value output by the dynamic change weight model to generate a decision vector, and perform cross-modal matching sorting based on the decision vector to output personalized content recommendation results.

[0012] Optionally, the extraction of the user purpose feature value, content contact feature value and time efficiency feature value in the expression vector, the construction of a dynamic change weight model combining the preference change trajectory of the historical behavior records, the integration of the user purpose feature value, content contact feature value, time efficiency feature value and change weight value output by the dynamic change weight model to generate a decision vector, and the cross-modal matching sorting based on the decision vector to output personalized content recommendation results, include:

[0013] Separate the user purpose feature value, content contact feature value and time efficiency feature value from the expression vector, and extract the operation time sequence based on the preference change trajectory of the historical behavior records, and extract the recent operation time and operation frequency from the operation time sequence;

[0014] Construct a dynamic change weight model based on the recent operation time and operation frequency and calculate a change weight value;

[0015] For each content item, calculate the user purpose feature value, the content contact feature value, the time efficiency feature value and the change weight value to generate a decision value corresponding to the content item, and organize all decision values into a decision vector according to the content item identifiers of different content items;

[0016] Uniformly sort the decision values of all content items in the decision vector in a cross-modal manner, and output the content item identifier with the highest sorting as the personalized recommendation result.

[0017] Optionally, for each content item, the user purpose feature value, the content contact feature value, and the time effectiveness feature value are respectively calculated with the change weight value to generate a decision value of the corresponding content item, and all decision values are organized into a decision vector according to content item identifiers of different content items, including:

[0018] extracting a user purpose feature value, a content contact feature value, and a time effectiveness feature value corresponding to the current content item;

[0019] multiplying the user purpose feature value by the change weight value to obtain a purpose weighted value, multiplying the content contact feature value by the change weight value to obtain a contact weighted value, and multiplying the time effectiveness feature value by the change weight value to obtain a time effectiveness weighted value;

[0020] adding the purpose weighted value, the contact weighted value, and the time effectiveness weighted value to generate a decision value of the current content item;

[0021] associating the decision value with a content item identifier of the current content item, and arranging all decision values in a decision vector according to the content item identifier after traversing all content items.

[0022] Optionally, the bipartite graph structure of the content node and the user node is constructed, the dynamic compensation coefficient is converted into an adjustment factor of connection strength to adjust the connection strength of the content node and the user node, and the associated weight of adjacent nodes is integrated to generate an expression vector containing association features, including:

[0023] constructing a bipartite graph structure composed of content nodes and user nodes, and assigning an initial connection strength value to a connection line between each content node and the user node based on the number of historical operations of the user on the content node;

[0024] converting the dynamic compensation coefficient into an adjustment factor of connection strength, and performing an operation on the initial connection strength value to obtain an adjusted connection strength value;

[0025] for each user node, obtaining the adjusted connection strength value and the associated weight of the adjacent content node, performing weighted processing on the connection strength value based on the associated weight, generating an association feature value of the corresponding user node, and arranging the association feature values of all user nodes in a node order to form an expression vector.

[0026] Optionally, the historical behavior records among different modal content in the content data stream are determined, and a dynamic compensation coefficient is generated by performing bias source analysis on the cross-modal conversion records in the historical behavior records and combining the similarity corresponding table, including:

[0027] reading a history record of behaviors among different modal contents in the content data stream from a storage system;

[0028] counting a proportion of a number of times of conversion from text content to picture content in the history record as a text-picture conversion value, and counting a proportion of a number of times of conversion from sound content to video content as a sound-video conversion value;

[0029] extracting a semantic relation score and a matching degree score in the similarity correspondence table, and calculating a first deviation amount of the text-picture conversion value and the semantic relation score and a second deviation amount of the sound-video conversion value and the matching degree score;

[0030] generating a dynamic compensation coefficient positively related to the deviation amount according to a size range of the first deviation amount and the second deviation amount.

[0031] Optionally, the analysis of semantic relations between text features and picture features and matching degrees between sound features and video contents among a plurality of independent feature packages by a pre-established large model to establish a similarity correspondence table across modalities comprises:

[0032] inputting text features in a text feature package and picture features in a picture feature package into the large model according to the independent feature packages, comparing content consistency of the text features and the picture features, and calculating a semantic relation score of each text feature and a corresponding picture feature;

[0033] inputting sound features in a sound feature package and video contents in a video content feature package into the large model, detecting a time alignment state of the sound features and the video contents, and calculating a matching degree score of each sound feature and a corresponding video content;

[0034] recording the semantic relation score as a text-picture relation item in a combination mode of the text feature package identifier and the picture feature package identifier, recording the matching degree score as a sound-video relation item in a combination mode of the sound feature package identifier and the video content feature package identifier, integrating all the text-picture relation items and the sound-video relation items, and constructing a similarity correspondence table containing identifier combinations and relation scores.

[0035] Optionally, the content data stream including multi-modal feature mixtures is obtained by configuring a multi-channel acquisition server, using a hardware separation unit of the multi-channel acquisition server to disassemble the content data stream into original data streams of different modal features, and generating independent feature packages according to corresponding feature data and identifiers of the original data streams, comprising:

[0036] obtaining a content data stream containing text, pictures, sound, and video;

[0037] The multi-channel acquisition server is configured to activate the plurality of physical channels of the hardware separation unit to split the content data stream, to generate original data streams including a text stream, a picture stream, a sound stream, and a video stream;

[0038] Character sequences and position information are extracted from the text stream, color distribution and contour point sets are extracted from the picture stream, frequency bands and loudness sequences are extracted from the sound stream, and motion trajectories and brightness changes are extracted from the video stream to extract feature data corresponding to each of the original data streams;

[0039] The feature data corresponding to the original data streams is appended with an identifier including a modal type code and a timestamp, and packaged to generate independent feature packages corresponding to text feature packages, picture feature packages, sound feature packages, and video feature packages.

[0040] In a second aspect, the application provides a personalized recommendation system based on a large model, comprising:

[0041] An acquisition module is configured to acquire a content data stream including a multi-modal feature mixture, split the content data stream into original data streams of different modal features by configuring a multi-channel acquisition server and using a hardware separation unit of the multi-channel acquisition server, and generate independent feature packages according to feature data and identifiers corresponding to the original data streams;

[0042] An analysis module is configured to analyze the semantic relationship between text features and picture features and the matching degree between sound features and video content among a plurality of the independent feature packages by a pre-established large model, to establish a similarity correspondence table across modalities;

[0043] A generation module is configured to determine historical behavior records among different modal content in the content data stream, and generate a dynamic compensation coefficient by combining a similarity correspondence table based on cross-modal conversion records in the historical behavior records;

[0044] An adjustment module is configured to construct a bipartite graph structure of content nodes and user nodes, convert the dynamic compensation coefficient into an adjustment factor of connection strength, adjust the connection strength of the content nodes and user nodes, and integrate the correlation weights of adjacent nodes to generate expression vectors including correlation features;

[0045] An output module is configured to extract user purpose feature values, content contact feature values, and time efficiency feature values in the expression vectors, construct a dynamic change weight model based on preference change trajectories of the historical behavior records, integrate the user purpose feature values, content contact feature values, time efficiency feature values, and change weight values output by the dynamic change weight model to generate a decision vector, and perform cross-modal matching sorting based on the decision vector to output a personalized content recommendation result.

[0046] Thirdly, this application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement a personalized recommendation method based on a large model as described in the first aspect above.

[0047] Fourthly, this application provides a computer storage medium storing a computer program, which, when executed by a computer, implements a personalized recommendation method based on a large model as described in the first aspect.

[0048] In this application example, a content data stream containing mixed multimodal features is acquired. A multi-channel acquisition server is configured, and its hardware separation unit decomposes the content data stream into raw data streams with different modal features. Independent feature packages are generated based on the feature data and identifiers corresponding to the raw data streams. A pre-established large model is used to analyze the semantic relationships between text and image features, and the matching degree between sound features and video content among multiple independent feature packages, to establish a cross-modal similarity correspondence table. Historical behavior records between different modalities in the content data stream are determined, and deviations are calculated based on the cross-modal conversion records in the historical behavior records and the similarity correspondence table. Source analysis generates dynamic compensation coefficients; a bipartite graph structure of content nodes and user nodes is constructed, and the dynamic compensation coefficients are transformed into connection strength adjustment factors to adjust the connection strength between content nodes and user nodes. The association weights of adjacent nodes are integrated to generate an expression vector containing association features. User purpose feature values, content connection feature values, and timeliness feature values ​​are extracted from the expression vector. A dynamic change weight model is constructed by combining the preference change trajectory of the historical behavior records. The user purpose feature values, content connection feature values, timeliness feature values, and change weight values ​​output by the dynamic change weight model are integrated to generate a decision vector. Based on the decision vector, cross-modal matching and ranking are performed to output personalized content recommendation results.

[0049] The technical solution of this application has the following beneficial effects:

[0050] This application efficiently decomposes multimodal content data streams into independent feature packages through multi-channel hardware separation. It utilizes a large model to establish a similarity correspondence table for cross-modal semantic relationships, analyzes the sources of deviation by combining cross-modal transformation records in historical behavior, and generates dynamic compensation coefficients. This drives the adaptive adjustment of the connection strength between content nodes and user nodes in the bipartite graph to construct an association expression vector. Finally, it integrates user purpose, content association, timeliness features, and the changing weights output by the dynamic weight model to generate a decision vector. This achieves high-precision matching and ranking of cross-modal content and user dynamic interests, significantly improving the accuracy and timeliness of personalized recommendations.

[0051] Further, after constructing the expression vector, user purpose feature values, content relevance feature values, and timeliness feature values ​​are separated from it. Simultaneously, operation time series are extracted based on historical behavior records, and the most recent operation time and frequency are analyzed. A dynamic weighting model is constructed using the most recent operation time and frequency, and variable weight values ​​are generated. For each content item, the user purpose feature value, content relevance feature value, and timeliness feature value are weighted and calculated with the variable weight value respectively to generate a content item-specific decision value, which is then organized into a decision vector according to content identifiers. Finally, the decision vectors are uniformly ranked across modalities, and the optimal content item identifier is output as the recommendation result. This scheme dynamically captures the user's most recent operation time and frequency to construct a weighting model, integrates multi-dimensional features with dynamic weights to generate content item-specific decision values, and achieves unified quantitative ranking of cross-modal content. This significantly improves the response speed and accuracy of recommendation results to users' instantaneous interests, while alleviating the problem of insufficient exposure of long-tail content.

[0052] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 A flowchart of a personalized recommendation method based on a large model provided in this application is shown;

[0055] Figure 2 The diagram illustrates a scenario of a personalized recommendation method based on a large model provided in this application.

[0056] Figure 3 A schematic diagram of the structure of a personalized recommendation system based on a large model provided in this application is shown;

[0057] Figure 4 A schematic diagram of the structure of a computing device provided in this application is shown. Detailed Implementation

[0058] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0059] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.

[0060] Research indicates that existing large-scale model recommendation solutions for cross-modal applications suffer from three limitations: first, reliance on offline-generated pseudo-data easily deviates from the true distribution, leading to distorted recommendations; second, static processing mechanisms cannot respond to changes in user behavior, causing recommendation lag when users' interests shift instantaneously; and third, shallow fusion of non-textual modalities results in cross-modal semantic fragmentation and insufficient value mining of multi-source information. These shortcomings stem from a lack of dynamic cross-modal alignment, interest tracking, and lightweight multimodal collaboration capabilities.

[0061] To address the aforementioned issues, this application proposes a personalized recommendation method based on a large model. This method generates independent feature packages by hardware-level multi-channel separation and decomposition of content data streams, and constructs a cross-modal semantic association table using a large model. It combines historical behavior to generate dynamic compensation coefficients, driving adaptive adjustment of the connection strength of bipartite graph nodes. Finally, it integrates user purpose, content association, timeliness features, and dynamic weights to generate a decision vector, achieving unified cross-modal ranking. This method eliminates data bias through hardware processing, bridges modal gaps through dynamic compensation, and captures both short-term and long-term interests through lightweight fusion, fundamentally solving the problems of recommendation distortion, lag, and modal fragmentation, significantly improving accuracy and timeliness.

[0062] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0063] Figure 1 A flowchart of a personalized recommendation method based on a large model is provided for embodiments of this application, such as... Figure 1 As shown, the method includes:

[0064] 101. Acquire a content data stream including multimodal feature mixture, configure a multi-channel acquisition server, use the hardware separation unit of the multi-channel acquisition server to decompose the content data stream into raw data streams with different modal features, and generate independent feature packages according to the feature data and identifiers corresponding to the raw data streams;

[0065] Optionally, step 101 may specifically include the following steps:

[0066] 1011. Obtain content data streams containing text, images, audio, and video;

[0067] 1012. By configuring a multi-channel acquisition server to activate multiple physical channels of the hardware separation unit, the content data stream is split to generate an original data stream containing text stream, image stream, audio stream, and video stream;

[0068] 1013. Extract character sequences and position information from the text stream, extract color distribution and contour point sets from the image stream, extract frequency bands and loudness sequences from the sound stream, and extract motion trajectories and brightness changes from the video stream, so as to extract feature data corresponding to each of the original data streams;

[0069] 1014. Add an identifier containing modality type code and timestamp to the feature data corresponding to the original data stream, and package them to generate independent feature packages corresponding to text feature packages, image feature packages, sound feature packages and video feature packages.

[0070] In the above scheme, the content data stream refers to the composite digital content carrier generated by user interaction, including synchronous or asynchronous combinations of text descriptions, static images, audio waveforms, and dynamic video images, which can be used to characterize the complete semantic expression of cross-media information. The multi-channel acquisition server refers to a computing device configured with dedicated hardware interfaces, which can be used to decompose the mixed data stream through physical isolation. The hardware separation unit refers to the physical signal decoupling module integrated into the server, which can be used to decompose the composite data stream into single-modal basic streams according to media type. The raw data stream refers to the single-modal basic data sequence generated by hardware separation, including plain text character streams, uncompressed image frame sequences, raw sound wave sampling sequences, or de-audio video frame sequences, which can be used to extract the essential features of a specific medium. Feature data refers to the set of quantitative indicators extracted from the raw data stream, including the positional sequence of text, the visual attributes of images, the physical parameters of sound waves, and the dynamic change trajectory of video, which can be used to characterize the essential attributes of each modality of content. Independent feature packages refer to structured and encapsulated feature datasets, containing feature data of a specific modality and its spatiotemporal identification metadata, which can be used to ensure data traceability and temporal consistency during cross-modal analysis.

[0071] In this embodiment of the application, the user-submitted content data stream containing text, images, audio, and video is first obtained through step 1011 at the platform data interface.

[0072] Subsequently, step 1012 activates four dedicated physical channels through the hardware separation unit of the multi-channel acquisition server to perform synchronous disassembly processing: the text channel uses a character signal filter based on regularity matching to remove non-text noise and outputs a pure character sequence stream; the image channel uses a trigger-based capture circuit based on inter-frame difference algorithm to identify key static frames and generate an image stream without dynamic redundancy; the audio channel uses a Fourier transform noise reduction chip to separate the target voiceprint and output a pure sound wave sampling sequence; the video channel uses a motion vector detector to filter still images and generate a motion stream without background interference. This process implements four-stream parallel processing at the hardware layer, ensures the integrity of each modality's data through physical isolation, and completes the disassembly of the 10-gigabit data stream within 80 milliseconds.

[0073] Then, starting from the raw data stream generated by hardware separation in step 1013, an OCR positioning algorithm is executed on the text stream to scan the character sequence and record the spatiotemporal coordinates, such as identifying the text "limited-time offer" at position [120, 80] at the 15th second of the video; edge detection and color histogram analysis are applied to the image stream to extract the body contour point set and the proportion of the main color tone, such as calculating that the red area occupies 65% in the product screenshot and marking 200 boundary points of the phone frame; a short-time Fourier transform is performed on the audio stream to decompose the frequency bands and draw the loudness change curve, such as separating the 300-500Hz human voice frequency band and recording the sudden increase in volume to 85dB at the "buy now" position; a dense optical flow tracking algorithm is used on the video stream to capture the motion trajectory and the brightness difference between frames, such as tracking the coordinate sequence of the demonstration gesture movement path and the light brightness jump value when the product is unpacked; then all feature data are normalized and encoded and output as a structured feature set to ensure the spatiotemporal reference consistency of cross-modal features.

[0074] Finally, in step 1014, each modal feature data is assigned a dual identifier: a modal type code and a millisecond-level timestamp. Text features are labeled "TXT", images are labeled "IMG", audio is labeled "AUD", and video is labeled "VID", with timestamps accurate to one-thousandth of a second, such as "20230815153025.456". Subsequently, structured encapsulation is performed, binding the character sequence and position information of the text stream into a text feature package, integrating the color distribution and contour point set of the image stream into an image feature package, packaging the frequency bands and loudness sequence of the audio stream into an audio feature package, and constructing the motion trajectory and brightness changes of the video stream into a video feature package. Using a binary serialization protocol, the data within each feature package is associated with and stored with its identifier, generating four standardized independent file packages. This ensures that cross-modal data can be aligned at the millisecond level via timestamps. For example, the motion trajectory data pointer in the video package points to the identifier "VID_20230815104533.456".

[0075] In practical applications, an online education platform receives a mixed data stream of "Chemical Experiment Operation Guide" uploaded by a user. This stream includes textual instructions for experimental steps, including the instruction "stop heating at 80°C," static photos of reagent bottles, a safety warning audio message about "wearing goggles," and a video of the solution's color change process. A multi-channel acquisition server activates a hardware separation unit. The text channel uses an ASCII filter to remove image noise and extract a pure text stream, outputting characters such as "stop heating at 80°C." The image channel uses a frame difference threshold algorithm, setting a close-up frame of the reagent bottle to generate an image stream when the difference between consecutive frames exceeds 5% pixels. The audio channel uses a 200-400Hz human voice bandpass filter to eliminate ambient noise and retain the pure audio stream of the "goggles" sound wave. The video channel uses motion vector detection to filter and fix the background, retaining the dynamic color change stream of the solution. Subsequently, feature extraction is performed. The text stream is spatiotemporally located using OCR to record the coordinates of "80°C" at the 120th second of the video (300, 150). The image stream is processed by H... SV color analysis statistics show that blue pixels account for 65% of the reagent bottle screenshot. The sound stream, using sound pressure level calculation, measures the peak volume of the word "goggles" to be 75dB, which is 60dB higher than the baseline volume. The video stream uses luminance difference to capture the instantaneous brightness jump during the solution's color change; the average brightness at frame 135 is 120 lumens, and at frame 136 it is 150 lumens, representing an improvement of 30 lumens. Finally, the data is packaged into four independent feature packets: a text packet binding instruction characters + spatiotemporal coordinates + the identifier "TXT_20". The image packet, labeled "230901103000.120", encapsulates 65% blue content data with the identifier "IMG_20230901103000.121". The audio packet integrates a 75dB peak value with the identifier "AUD_20230901103000.122". The video packet is associated with a 30-stream brightness difference with the identifier "VID_20230901103000.123". These four independent feature packets have millisecond-level continuous timestamps to ensure accurate cross-modal alignment.

[0076] The above-mentioned 101 overall solution transforms complex mixed data into standardized independent feature packages through hardware-level decomposition and feature extraction, providing structured, high-quality input for subsequent cross-modal analysis, while ensuring accurate alignment and traceability of different types of data.

[0077] 102. By using a pre-established large model, analyze the semantic relationship between text features and image features, and the degree of matching between sound features and video content among multiple independent feature packages, in order to establish a cross-modal similarity correspondence table;

[0078] Optionally, step 102 may specifically include the following steps:

[0079] 1021. Based on the independent feature package, input the text features in the text feature package and the image features in the image feature package into the large model, compare the content consistency between the text features and the image features, and calculate the semantic relationship score between each text feature and the corresponding image feature;

[0080] 1022. Input the sound features in the sound feature package and the video content in the video content feature package into the large model, detect the time alignment status of the sound features and the video content, and calculate the matching degree score between each sound feature and the corresponding video content.

[0081] 1023. Record the semantic relationship score as a text-image relationship item according to the combination of the text feature package identifier and the image feature package identifier, and record the matching degree score as an audio-video relationship item according to the combination of the sound feature package identifier and the video content feature package identifier. Integrate all the text-image relationship items and the audio-video relationship items to construct a similarity correspondence table containing identifier combinations and relationship scores.

[0082] In the above scheme, the semantic relationship score is a quantitative indicator reflecting the consistency between text descriptions and image content. It includes matching strength information between text semantic vectors and visual feature vectors, and can be used to evaluate the semantic alignment quality between cross-modal content. The matching degree score is a measure characterizing the spatiotemporal synchronization between sound features and video content. It includes the matching accuracy information between audio event timestamps and video action frame sequences, and can be used to detect the authenticity of audio-visual collaborative performance. The similarity correspondence table is a global mapping database that integrates the correlations between multimodal features. It contains the spatial topological structure information of all text-image relationship items and audio-video relationship items, and can be used to drive the decision engine for cross-modal content collaborative analysis.

[0083] In this embodiment, firstly, step 1021 inputs the text semantic vector from the text feature package and the visual vector from the corresponding image feature package into a pre-trained large model. The model aligns the text and image regions segment by segment using a cross-modal attention mechanism: extracting key concepts from the text and calculating the semantic overlap between each text segment and the relevant image region. Then, a cosine similarity weighted algorithm is used to synthesize the alignment results of all segments to generate an overall semantic relationship score. For example, when the text describes "red surfboard", the model detects the color features of the surfboard in the image. If the color value of the region matches the red spectrum range, such as RGB(220,20,60), a high score is assigned, and the final output semantic relationship score of the text-image combination is 0.93.

[0084] Secondly, in step 1022, the sound feature package and video content package processed in step 1021 are input into the same large model. The model aligns the timeline using a dynamic time warping algorithm, extracts the timestamps of audio events and video action frames, and calculates the overlap ratio of their time windows. Simultaneously, an audio-visual consistency detection module verifies content matching, such as whether laughter corresponds to a person opening their mouth. Finally, a matching score is generated based on time deviation and content consistency. For example, if the video shows a person raising a glass in the frame sequence from 5.1 to 5.4 seconds, and the audio shows a glass-clattering sound at 5.15 seconds, the model determines that the time deviation is <0.2 seconds and the sound source is consistent, outputting a matching score of 0.88.

[0085] Finally, step 1023 receives the semantic relationship score generated in step 1021 and its corresponding text feature package identifier and image feature package identifier, binding the three into a structured relationship unit. For example, for the product description text package identifier T07 and the main image package identifier P12 on an e-commerce platform, based on the semantic relationship score of 0.93 calculated in step 1021, a "text-image" relationship item is generated, fully recorded as "identifier combination T07_P12, semantic relationship score 0.93". Simultaneously, the matching degree score output in step 1022 and its associated audio feature package identifier and video content package identifier are received. For example, the matching score of the product explanation audio package identifier S09 and the demonstration video package identifier V15 is 0.88, generating the "audio-video" relationship item "identifier combination S09_V15, matching degree score 0.88". Finally, the data aggregation engine integrates all the above relationships in a unified format: assigning a unique index key to each relationship and building a globally queryable similarity correspondence table. For example, text and image items are stored with the prefix "TP_", and audio and video items are stored with the prefix "SV_".

[0086] In practical applications, when a short video platform recommends science content to users, the system acquires a mixed data stream of wildlife documentaries. The text feature package contains the semantic vector of the narration "the cheetah's spinal extension and contraction provides explosive power during acceleration," and the corresponding image feature package contains the visual vectors of a burst of footage of a cheetah running. The large model decomposes the text into three segments: "cheetah body," "spine extension and contraction," and "explosive power," respectively, and matches them to the cheetah outline region, slow-motion spinal deformation frames, and dust effect region in the image. The matching scores are 0.97, 0.88, and 0.91, respectively, and the weighted average semantic relationship score is 0.92. Subsequently, the audio and video of the same documentary are processed: the sound feature package contains the spectrum of the cheetah's panting and the sound of its paws stomping the ground, and the video content package contains 20 seconds of keyframes of the cheetah sprinting. The model detected a 0.02-second time difference between the image of the cheetah's claws striking the ground at 5.3 seconds in the video and the impact sound at 5.28 seconds in the audio, and a 0.01-second time difference between the panting image at 8.1 seconds and the breathing sound at 8.09 seconds. Taking the average difference of 0.015 seconds, the matching formula yielded a matching score of 0.87. Finally, the semantic relationship score (0.92) between the text and image feature packages was bound as one relation, and the matching score (0.87) between the sound and video feature packages was bound as another relation, integrated into a similarity table. When a user browses the "Animal Movement Mechanisms" graphic and text content, the system uses this table to associate highly matching video clips, triggering personalized recommendations of cheetah sprint videos.

[0087] The overall solution described above (102) utilizes a large-scale model to deeply analyze the semantic consistency between text and images, and the spatiotemporal synchronization between sound and video, transforming the complex relationships between cross-modal content into a quantifiable similarity table. This table accurately characterizes the matching strength between multimodal features, providing crucial information for subsequent dynamic compensation of cross-modal biases. It significantly enhances the system's ability to understand text-image relationships and audio-visual synergy, laying a solid foundation for accurate matching of multi-source heterogeneous content in personalized recommendations, and ultimately driving the simultaneous optimization of recommendation accuracy and timeliness.

[0088] 103. Determine the historical behavior records between different modal content in the content data stream, and generate dynamic compensation coefficients based on the cross-modal conversion records in the historical behavior records and the similarity correspondence table to analyze the source of deviation.

[0089] Optionally, step 103 may specifically include the following steps:

[0090] 1031. Read historical behavior records between different modalities of content in the content data stream from the storage system;

[0091] 1032. Calculate the proportion of times text content is converted to image content in the historical behavior records as the text-to-image conversion value, and at the same time calculate the proportion of times audio content is converted to video content as the audio-to-video conversion value;

[0092] 1033. Extract the semantic relationship score and matching degree score from the similarity correspondence table, and calculate the first deviation between the text image conversion value and the semantic relationship score and the second deviation between the audio video conversion value and the matching degree score;

[0093] 1034. Based on the magnitude range of the first deviation and the second deviation, generate a dynamic compensation coefficient that is positively correlated with the deviation.

[0094] In the above scheme, historical behavior records refer to the user's interaction logs with different modalities of content such as text, images, audio, and video, including browsing duration, click frequency, and content switching behavior, used to reconstruct the user's content consumption habits. Cross-modal conversion records specifically refer to the statistical proportion of users switching from one content modality to another, reflecting the user's cross-modal interest migration tendency. The similarity mapping table is a correlation strength mapping table generated by the large model, recording the semantic relevance scores of text and images, and the content matching scores of audio and video, serving as a benchmark reference for cross-modal association relationships. Deviation source analysis locates the direction of deviation in the system's understanding of user interests by comparing the difference between the user's actual switching proportion and the association scores predicted by the large model. The dynamic compensation coefficient is a correction weight value calculated based on the deviation amount; the larger the deviation, the higher the coefficient, used to enhance underestimated cross-modal association relationships in subsequent recommendations.

[0095] In this embodiment, step 1031 first extracts the target user's historical behavior logs from the storage database. These logs record user interaction events with different modal content such as text, images, audio, and video in chronological order. Key fields, including content identifier, modality type, operation type, and timestamp, are filtered using a structured query language to form an operation sequence list sorted by time. For example, user A generated 150 behavior records in the past week, with the 5th record being "clicked on text report ID123" and the 32nd record being "switched from text report ID123 to image set ID456".

[0096] Next, based on the operation sequence list generated in step 1031, step 1032 iterates through all adjacent operation records to identify nodes where the modality type changes. For text-to-image conversion scenarios, the number of times "text content is immediately followed by an operation on image content" is counted, and then divided by the total number of cross-modal switches to calculate the text-to-image conversion value. For example, if user A has 150 records with 20 such conversions and a total of 50 cross-modal switches, then the text-to-image conversion value is... The audio-to-video conversion value is calculated synchronously as described above. For example, if user A has 150 records with 10 such conversions, and the total number of cross-modal switching is 50, then the audio-to-video conversion value is... The final output includes two key metrics: text-to-image conversion value and audio-to-video conversion value.

[0097] Then, in step 1033, the similarity correspondence table of cross-modal association strength analyzed by the pre-generated large-scale storage model is invoked. Semantic relationship scores for text and images and matching scores for audio and video are extracted. The output value of step 1032 is compared with the corresponding scores to calculate the first and second deviations, using the following formula: , For example, when user B's text-image conversion value is 0.3, while the semantic relationship score of text-image in the similarity table is 0.8, the system calculates... As the first deviation, the absolute difference between its audio-video conversion value of 0.2 and its matching score of 0.5 is calculated synchronously. As the second deviation measure, these two deviation measures directly quantify the error direction and magnitude of the system's understanding of the user's cross-modal interests, providing a basis for subsequent dynamic compensation.

[0098] Finally, step 1034 receives the first and second deviations output in step 1033, adds them together to obtain the total deviation, and then generates a dynamic compensation coefficient according to a preset compensation rule mapping table. This rule stipulates that the larger the total deviation, the higher the compensation coefficient. Specifically, this is achieved through a piecewise function: when the total deviation is less than 0.4, a linear function is used. When the total deviation is in the range of 0.4 to 0.6, the following method is used: When the total deviation exceeds 0.6, the enhanced compensation function is activated. This process ensures that the compensation coefficients remain within the range of 0.3 to 0.9 and are strictly positively correlated with the degree of deviation between the user's actual cross-modal behavior and the system's prediction. The final generated coefficients will directly affect the adjustment ratio of the subsequent bipartite graph connection strength. For example, if user C's text-image deviation of 0.4 and audio-visual deviation of 0.3 are added together to form a total deviation of 0.7, which is greater than 0.6, then this deviation is substituted into the reinforcement compensation function for calculation. This means that the system will increase the weight of underestimated cross-modal content association by 78% in subsequent recommendations.

[0099] In practical applications, within the knowledge-sharing platform E, the system reads 200 behavioral records from user F's stored logs over the past 30 days, containing 80 cross-modal switching operations. By traversing the operation sequences, the specific conversion behaviors are identified: Statistical analysis reveals 32 text-to-image switches (e.g., switching from a programming tutorial text to a code illustration), and 16 audio-to-video switches (e.g., switching from an algorithm explanation audio to a demonstration video). Therefore, the text-to-image conversion value is calculated as 32 divided by the total number of switches (80), resulting in a text-to-image conversion value of 0.4. The audio-to-video conversion value is calculated as 16 divided by 80, resulting in an audio-to-video conversion value of 0. 2; Next, the platform's pre-stored similarity correspondence table is called to extract the semantic relationship score of text and image (0.75) and the matching degree score of audio and video (0.55). The first deviation is calculated as the absolute value of 0.4 minus 0.75, which equals 0.35, and the second deviation is calculated as the absolute value of 0.2 minus 0.55, which also equals 0.35. Finally, the two deviations are added together to obtain the total deviation of 0.7. According to the preset rule "when the total deviation is ≥ 0.6, the compensation coefficient = 0.4 × total deviation + 0.5", the formula is 0.4 × 0.7 + 0.5 = 0.78, generating a dynamic compensation coefficient of 0.78 to correct the subsequent content recommendation weight for user F.

[0100] The aforementioned overall solution (103) dynamically generates compensation coefficients by quantifying the difference between the actual cross-modal switching behavior of users and the correlation strength predicted by the large model. This corrects the recommendation system's misunderstanding of the correlation between text, images, audio, and video content. Based on a compensation mechanism driven by real behavioral data, it eliminates recommendation distortion caused by static model predictions, responds instantly to shifts in user interests, and bridges semantic barriers between different modalities, significantly improving the accuracy and timeliness of matching cross-modal content with users' dynamic needs.

[0101] 104. Construct a bipartite graph structure of content nodes and user nodes, convert the dynamic compensation coefficient into an adjustment factor for connection strength, adjust the connection strength between content nodes and user nodes, and integrate the association weights of adjacent nodes to generate an expression vector containing association features.

[0102] Optionally, step 104 may specifically include the following steps:

[0103] 1041. Construct a bipartite graph structure consisting of content nodes and user nodes, and assign an initial connection strength value to the connection line between each content node and the user node based on the number of historical operations performed by the user on the content node;

[0104] 1042. Convert the dynamic compensation coefficient into an adjustment factor for the connection strength, and perform a calculation on the initial connection strength value to obtain the adjusted connection strength value;

[0105] 1043. For each user node, obtain the adjusted connection strength value and association weight of its adjacent content nodes, perform weighted processing on the connection strength value based on the association weight, generate the association feature value of the corresponding user node, and arrange the association feature values ​​of all user nodes in node order to form an expression vector.

[0106] In the above scheme, the dynamic compensation coefficient refers to a quantitative value reflecting the need for cross-modal data deviation correction. It is derived from the deviation analysis results of cross-modal conversion records and similarity correspondence tables in user history behavior and is used to adjust the association strength between user nodes and content nodes. The expression vector refers to a structured numerical sequence representing the overall distribution of user interests, composed of the sequentially arranged association feature values ​​of all user nodes, used to drive the unified ranking decision of cross-modal content. The association feature value refers to the quantitative interest index of a single user node, generated by weighted aggregation of the adjusted connection strength of its adjacent content nodes, used to construct a dynamic interest profile at the user level. The adjustment factor refers to the dynamic scaling coefficient of the connection strength, directly transformed from the dynamic compensation coefficient, containing information on the direction and magnitude of cross-modal deviation correction, used to adaptively update the association relationship between users and content in the bipartite graph. The association weight refers to the inherent importance parameter of the content node, statically set based on content attributes, including a popularity decay coefficient and a long-tail value gain factor, used to balance the contribution ratio of popular and unpopular content in interest representation.

[0107] In this embodiment, firstly, step 1041 maps all platform users to independent user nodes, and simultaneously maps all content items to independent content nodes, forming an initial set of two types of nodes. When a user has performed a historical operation on a content item, the system establishes a connection between the corresponding user node and the content node. Next, the total number of historical operations performed by the user on that content item corresponding to each connection is counted, and this number is multiplied by a preset base weight coefficient of 0.2 to obtain the initial connection strength value. Finally, a bipartite graph network structure containing user nodes, content nodes, and connection lines with strength labels is constructed.

[0108] Next, in step 1042, the dynamic compensation coefficient generated in step 103 is received as an adjustment factor, and the connections between all user nodes and content nodes in the bipartite graph are traversed. For each connection, the initial connection strength value calculated in step 1041 is read, and this value is multiplied by the adjustment factor for scaling to generate an updated connection strength value, which replaces the original value. This process completes the correction of the correlation strength between all network nodes, enabling the bipartite graph relationships to dynamically adapt to changes in user interests. For example, when user U1's cross-modal behavior triggers a dynamic compensation coefficient of 1.1, the system traverses its connection lines: the original U1-V1 line strength of 0.6 is multiplied by 1.1 to obtain a new strength of 0.66, and the original U1-V2 line strength of 0.2 is multiplied by 1.1 to obtain a new strength of 0.22. The updated bipartite graph reflects the user's increased interest in video content.

[0109] Finally, in step 1043, each user node is traversed to find all content nodes directly connected to it as neighboring nodes. For each neighboring node, the updated connection strength value from step 1042 and the predefined association weight value for that content are obtained, and the two are multiplied to obtain a weighted strength value. The weighted strength values ​​of all neighboring nodes of the user are accumulated to generate the association feature value of the user. Finally, the association feature values ​​of all users are arranged into a numerical sequence according to user number, forming an expression vector and output to the recommendation decision module. For example, the neighboring nodes of user U1 are videos V1 and V2, where the adjusted strength and weight of V1 are 0.66 and 0.4, respectively, and the adjusted strength and weight of V2 are 0.22 and 0.6, respectively; the weighted accumulated value is calculated as follows: 0.66 multiplied by 0.4 equals 0.264, 0.22 multiplied by 0.6 equals 0.132, and the two are added together to obtain the association feature value of 0.396. When the three user feature values ​​in the system are 0.396, 0.51, and 0.23 respectively, the final output expression vector is [0.396, 0.51, 0.23].

[0110] In practical applications, on a music streaming platform, user U1 favorites song S1 twice and reads an interview article with singer P1 five times. The system maps U1 as a user node, and song S1 and article P1 as content nodes, establishing two connections: one from U1 to S1 and the other from U1 to P1. The initial connection strength is calculated based on the historical number of operations: the U1-S1 connection strength is 2 multiplied by the base weight coefficient 0.2, resulting in 0.4; the U1-P1 connection strength is 5 multiplied by 0.2, resulting in 1.0. Because the user recently shifted from listening to music to reading, a dynamic compensation coefficient of 0.9 is applied. The initial strength of each connection is then multiplied by this coefficient to update the value: the new strength of U1-S1 is 0.4 multiplied by 0.9, resulting in 0.36; the new strength of U1-P1 is 1.0 multiplied by 0.9, resulting in 0.9. Then, the neighboring nodes S1 and P1 of U1 are obtained. S1 corresponds to a weight of 0.36 for new intensity and 0.5 for song category, and P1 corresponds to a weight of 0.9 for new intensity and 0.7 for song category. The weighted values ​​0.36 multiplied by 0.5 equal 0.18 and 0.9 multiplied by 0.7 equal 0.63, respectively. They are added together to obtain the associated feature value 0.81. Finally, the feature value of U1 0.81, the feature value of user U2 0.65, and the feature value of U3 0.42 are arranged in ascending order by user ID to generate the expression vector [0.81, 0.65, 0.42], which is then input into the recommendation module.

[0111] The overall solution described above (104) constructs a user-content bipartite graph network, initializes the connection strength between nodes based on historical behavior, and establishes a basic association model. It then uses dynamic compensation coefficients to adjust the connection strength values, enabling the network relationships to adaptively respond to user interest shifts. Finally, it weightedly fuses information from adjacent nodes to generate an expression vector, achieving quantitative representation and dynamic updating of user interests. This process effectively solves the problems of static networks failing to capture interest drift and fragmented cross-modal associations, providing a newer, structured data foundation for accurate recommendations.

[0112] 105. Extract the user purpose feature value, content connection feature value, and timeliness feature value from the expression vector, construct a dynamic change weight model by combining the preference change trajectory of the historical behavior record, integrate the user purpose feature value, content connection feature value, timeliness feature value and the change weight value output by the dynamic change weight model to generate a decision vector, and output personalized content recommendation results based on the decision vector by performing cross-modal matching and ranking.

[0113] Optionally, step 105 may specifically include the following steps:

[0114] 1051. Separate user purpose feature value, content connection feature value and timeliness feature value from the expression vector, and extract operation time series based on the preference change trajectory of historical behavior records, and extract the most recent operation time and operation frequency from the operation time series;

[0115] 1052. Construct a dynamic weighting model based on the most recent operation time and operation frequency, and calculate the weighting values.

[0116] 1053. For each content item, the user purpose feature value, the content connection feature value, and the timeliness feature value are calculated with the change weight value to generate the corresponding decision value for the content item, and all decision values ​​are organized into a decision vector according to the content item identifier corresponding to different content items.

[0117] Step 1053 may specifically include the following processes: extracting the user purpose feature value, content connection feature value, and timeliness feature value corresponding to the current content item; multiplying the user purpose feature value by a change weight value to obtain a purpose weighted value, multiplying the content connection feature value by a change weight value to obtain a connection weighted value, and multiplying the timeliness feature value by a change weight value to obtain a timeliness weighted value; adding the purpose weighted value, the connection weighted value, and the timeliness weighted value to generate the decision value of the current content item; associating the decision value with the content item identifier of the current content item, and after traversing all content items, arranging all the decision values ​​in the order of the content item identifier to form a decision vector.

[0118] 1054. Perform cross-modal unified sorting on the decision values ​​of all content items in the decision vector, and output the identifier of the content item with the highest ranking as the personalized recommendation result.

[0119] In the above scheme, the user purpose feature value refers to a quantitative indicator reflecting the user's current core intent, including behavioral data such as search keyword strength and page dwell time, which can be used to identify the user's dominant demand direction. The content relevance feature value refers to a metric representing the correlation between content items and the user's historical preferences, including multi-dimensional information such as topic similarity, author attention, and interaction depth, which can be used to assess the degree of matching between content and user interests. The timeliness feature value refers to a dynamic value measuring the freshness of content and its relevance to recent user behavior, including time-sensitive factors such as publication time decay coefficient and interaction frequency, which can be used to capture interest drift trends. The dynamically changing weight model refers to a weight calculation function built based on user operation behavior, including a synergistic mechanism of recent operation time weight factor and operation frequency weight factor, which can be used to dynamically adjust feature importance. The decision vector refers to a recommendation priority sequence formed by integrating weighted feature values, including the fusion decision value of all content items and their unique identifier, which can be used to achieve fair ranking output of cross-modal content.

[0120] In this embodiment, the axis is separated from the expression vector in step 1051. This value is calculated by analyzing behavioral data such as the strength of user search keywords and page dwell time, and is used to quantify the user's current core intent direction. At the same time, the content relevance feature value is separated, and the similarity between the content item and the user's historical preferences and the author's attention are evaluated based on the topic matching algorithm. Then, the timeliness feature value is separated, and the freshness weight is dynamically calculated by combining the time decay function with the content publication time and interaction frequency. Meanwhile, the system retrieves the operation time sequence in the user's historical behavior record, that is, the complete interaction log sorted by timestamp, and parses out the specific time point of the most recent operation and the operation frequency statistics within the set time window.

[0121] Next, in step 1052, based on the recent operation time interval and operation frequency data provided in step 1051, the dynamic weight calculation engine is activated for processing. For the recent operation time, an inverse proportional function is used to calculate the time weight factor, the core principle being that the smaller the time interval, the higher the weight; for example, a 2-hour interval generates a coefficient of 1.2. For the operation frequency, a logarithmic function is used to calculate the frequency weight factor, its characteristic being that the weight increase gradually converges as the frequency increases; for example, 7 operations within 24 hours generate a coefficient of 1.4. The time weight factor and the frequency weight factor are then multiplied to obtain the comprehensive change weight value. For example, when it is detected that the user's most recent operation occurred 1 hour ago and the user has clicked on the sports shoe category 12 times that day, the time weight factor is calculated as 1.3 according to the inverse proportional principle, and the frequency weight factor is calculated as 1.5 according to the logarithmic function characteristic, ultimately generating a change weight value of 1.95.

[0122] Then, in step 1053, for each content item to be recommended, the corresponding user purpose feature value, content connection feature value, and timeliness feature value are extracted. These three feature values ​​are then multiplied by the change weight value generated in step 1052 to obtain the purpose weight value, connection weight value, and timeliness weight value. Next, the three weight values ​​are summed to calculate the final decision value for the content item. Finally, the unique identifier of the content item is associated with the decision value. After traversing all candidate content items, all decision values ​​are integrated in identifier order to form a decision vector. For example, for running shoes, a product frequently viewed by users, the user purpose feature value (0.9), content connection feature value (0.7), and timeliness feature value (0.8) are extracted and weighted by a change weight value of 1.6 to obtain the final decision value. , , The sum of the three values ​​gives a decision value of 3.84. After associating the product identifier ID_RUN001, this decision value is entered into the vector and eventually occupies the first place in the decision vector containing 20 products.

[0123] Finally, after receiving the decision vector in step 1054, a cross-modal unified ranking mechanism is initiated to globally compare the decision values ​​of all content items. This mechanism uses an extreme value retrieval algorithm to scan the entire decision vector, automatically locating the decision value with the highest value and locking its corresponding unique identifier for the content item. The content item represented by this identifier is then directly output as the final personalized recommendation result. For example, when the decision value of the hiking pole video (4.37) in the decision vector surpasses that of the text and image guide (3.12), the system automatically pushes the video to the user's homepage, achieving optimal matching in a cross-modal scenario.

[0124] In practical application, after user D searches for "family camping equipment" on the comprehensive content platform, the system first extracts three key values ​​from their behavioral expression vector: a user purpose feature value of 0.75 reflecting a strong outdoor shopping intention; a content relevance feature value of 0.65 indicating the relevance to previously purchased tent products; and a timeliness feature value of 0.85 corresponding to newly listed portable grill products. Simultaneously, the system detects that the most recent operation occurred 3 hours ago, and that the user has browsed the camping category 6 times within the past 24 hours. The system then initiates dynamic weight calculation: based on a 3-hour interval, a time factor of 1.15 is generated using an inverse proportional function; combined with the frequency of the 6 operations, a frequency factor of 1.45 is generated using a logarithmic function; multiplying these two factors yields a comprehensive change weight value of 1.15 multiplied by 1.45, which equals 1.6675. Next, for the portable grill product: the target feature value (0.75) is multiplied by its weight value to get 1.250625, the content relevance feature value (0.65) is multiplied by its weight value to get 1.083875, and the timeliness feature value (0.85) is multiplied by its weight value to get 1.417375. These three are added together to generate the final decision value: 1.250625 + 1.083875 + 1.417375 = 3.751875, which is then associated with the product identifier ID_Grill202. After the system has traversed 32 candidate products, including picnic mats and folding chairs, the portable grill, with a decision value of 3.751875, surpasses the picnic mat's 2.98 and the folding chair's 3.21, ultimately winning in the cross-modal ranking. The product video is then pushed to the user's homepage. At this time, the user is planning a weekend family camping trip, and this recommendation accurately matches their immediate purchasing needs.

[0125] The aforementioned overall solution (105) separates user intent, content relevance, and timeliness features, and combines recent behavior frequency with dynamic time-based weighting models to achieve precise weighted fusion of multi-dimensional features. Based on the weighted results, it generates content-specific decision values ​​and forms globally comparable vectors, ultimately outputting the optimal recommendation results through cross-modal unified ranking. This mechanism effectively captures users' instantaneous changes in interest, synchronously coordinates long-term preferences and needs, overcomes recommendation biases caused by static data processing and modal barriers in traditional solutions, significantly improves the timeliness and accuracy of content matching, and ensures that high-frequency, high-demand content receives priority exposure.

[0126] The following is a complete example for steps 101-105, such as Figure 2 As shown, when user Li Hua searches for high-altitude camping equipment on an outdoor equipment platform, the multi-channel acquisition server captures a mixed data stream containing a high-altitude tent selection guide with mixed text and images, a user-uploaded gas stove usage review video, and a voice Q&A session about high-altitude cooking utensils. The hardware separation unit uses a dedicated image processing chip to split the text and images into text feature streams and image feature streams. The text stream extracts text features such as titles and parameter descriptions, while the image stream analyzes tent structure diagrams and material details. Simultaneously, the audio decoder separates the video's frame sequences from the audio stream. The frame sequences capture the gas stove operation demonstration, while the audio stream separates the narration. Finally, five independent feature packages are generated and labeled with a unified content identifier: the text feature package contains keywords such as wind resistance and warmth retention; the image feature package stores tent wind resistance test diagrams; the video feature package records the gas stove ignition demonstration; and the audio feature package stores explanations of key usage points at high altitudes. All feature packages achieve cross-modal association through content identifiers.

[0127] Subsequently, the platform invokes a pre-trained multimodal large model to analyze the correlation between feature packets. The model first calculates the semantic consistency between the wind resistance description in the text packet and the tent deformation test image in the image packet, obtaining a similarity score of 0.91 through an attention matching algorithm. Then, it detects the spatiotemporal alignment between the close-up frame of the blue flame in the video packet and the narration of insufficient oxygen leading to incomplete combustion in the audio packet, identifying a 0.8-second delay between the keyframe and the audio and giving a matching value of 0.87. Based on this, a cross-modal similarity correspondence table is constructed, establishing mapping entries from the wind resistance parameters of the text and images to the test image, while simultaneously generating association entries from the abnormal flame phenomena in the video and audio to the explanation of incomplete combustion.

[0128] The system reviewed user Li Hua's behavior over the past 30 days and found that he frequently clicked on tent images and text, then immediately watched test videos. The conversion rate from images and text to videos reached 89%, but the average completion rate for the gas stove videos was only 45%. A deep analysis using a similarity table revealed the following sources of discrepancies: the cross-modal matching degree between the tent images / text and videos reached 0.91, meeting expectations and requiring no compensation; the gas stove content suffered from technical discrepancies due to a lack of synchronization between the visuals and narration, as well as content defects including a missing high-altitude scene. Based on this, a dynamic compensation coefficient was generated: the tent content maintained a baseline coefficient of 1.0, while the gas stove content received additional compensation of 0.1 for technical discrepancies and 0.15 for missing content, ultimately determining the compensation coefficient to be 1.25.

[0129] Then, a bipartite graph structure is constructed between the user Li Hua and the content nodes, with the user node connected to the tent (text / image) node and the gas stove (video) node. The initial edge strength is set to 0.85 for the tent connection and 0.75 for the gas stove connection. A compensation coefficient is applied to dynamically adjust the connection relationships: the gas stove edge strength is updated to 0.75 multiplied by 1.25, resulting in 0.9375. Features of adjacent nodes are aggregated through a graph convolutional network, integrating the material parameters of the tent node, the high-altitude suitability of the gas stove node, and the cold-resistance requirements from the user's historical behavior. Finally, an expression vector is generated containing a professional requirement value of 0.83, a scene adaptability value of 0.78, and a timeliness sensitivity value of 0.95.

[0130] Finally, from the expression vector, the user's purpose feature value (0.83) represents the intention to purchase high-altitude equipment, the content-related feature value (0.78) corresponds to the gas stove, and the timeliness feature value (0.95) corresponds to the newly listed high-altitude-specific stove. User dynamics were monitored synchronously: four searches for high-altitude stoves were conducted within the last fifteen minutes. The dynamic weight model calculated the time factor and frequency factor: an inverse proportional function was used to process the fifteen-minute interval, outputting 1.4; a logarithmic function was used to process the four-time frequency, outputting 1.35; the comprehensive weight value was 1.4 multiplied by 1.35, equaling 1.89. The gas stove decision value was calculated step-by-step: the purpose weighted value (0.83) multiplied by 1.89 yielded 1.5687; the content-related weighted value (0.78) multiplied by 1.89 yielded 1.4742; and the timeliness weighted value (0.95) multiplied by 1.89 yielded 1.7955. The sum of these three values ​​yielded a decision value of 4.8384. When this value exceeds the tent decision value of 4.12, the system will push a video on the homepage explaining the operation of the plateau gas stove with a newly added real-shot frame of ignition at an altitude of 5000 meters while the user is packing their luggage for the trip to Tibet.

[0131] Figure 3 This application provides a schematic diagram of the structure of a personalized recommendation system based on a large model, as shown in the embodiments below. Figure 3 As shown, the system includes:

[0132] The acquisition module 31 is used to acquire a content data stream including a mixture of multimodal features. By configuring a multi-channel acquisition server, the hardware separation unit of the multi-channel acquisition server is used to decompose the content data stream into raw data streams with different modal features, and generate independent feature packages based on the feature data and identifiers corresponding to the raw data streams.

[0133] Analysis module 32 is used to analyze the semantic relationship between text features and image features, and the degree of matching between sound features and video content among multiple independent feature packages through a pre-established large model, so as to establish a cross-modal similarity correspondence table;

[0134] The generation module 33 is used to determine the historical behavior records between different modal content in the content data stream, and to generate dynamic compensation coefficients based on the cross-modal conversion records in the historical behavior records and the similarity correspondence table to perform deviation source analysis.

[0135] The adjustment module 34 is used to construct a bipartite graph structure of content nodes and user nodes, convert the dynamic compensation coefficient into an adjustment factor of connection strength, so as to adjust the connection strength of the content nodes and user nodes, and integrate the association weights of adjacent nodes to generate an expression vector containing association features.

[0136] Output module 35 is used to extract user purpose feature value, content connection feature value and timeliness feature value from the expression vector, construct a dynamic change weight model by combining the preference change trajectory of the historical behavior record, integrate the user purpose feature value, content connection feature value, timeliness feature value and the change weight value output by the dynamic change weight model to generate a decision vector, and perform cross-modal matching and ranking based on the decision vector to output personalized content recommendation results.

[0137] Figure 3 The aforementioned personalized recommendation system based on a large model can execute... Figure 1 The implementation principle and technical effects of the large-model-based personalized recommendation method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit performs operations in the large-model-based personalized recommendation system described in the above embodiments have been detailed in the embodiments related to this method, and will not be elaborated upon here.

[0138] In one possible design, Figure 3 The personalized recommendation system based on a large model shown in the embodiment can be implemented as a computing device, such as... Figure 4 As shown, the computing device may include a storage component 41 and a processing component 42;

[0139] The storage component 41 stores one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component 42.

[0140] The processing component 42 is used for the above Figure 1 The embodiment describes a personalized recommendation method based on a large model.

[0141] The processing component 42 may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0142] Storage component 41 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0143] Of course, computing devices may also include other components, such as input / output interfaces, display components, communication components, etc.

[0144] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.

[0145] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.

[0146] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.

[0147] This application also provides a computer storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 1 The illustrated embodiment presents a personalized recommendation method based on a large model.

[0148] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0149] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0150] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A personalized recommendation method based on a large model, characterized in that, include: The content data stream, which includes a mixture of multimodal features, is acquired. By configuring a multi-channel acquisition server, the hardware separation unit of the multi-channel acquisition server is used to decompose the content data stream into raw data streams with different modal features, and independent feature packages are generated based on the feature data and identifiers corresponding to the raw data streams. By using a pre-established large model, the semantic relationship between text features and image features, and the degree of matching between sound features and video content among multiple independent feature packages are analyzed to establish a cross-modal similarity correspondence table. Determine the historical behavior records between different modal content in the content data stream, and based on the cross-modal conversion records in the historical behavior records, combine the similarity correspondence table to perform deviation source analysis and generate dynamic compensation coefficients; A bipartite graph structure of content nodes and user nodes is constructed, and the dynamic compensation coefficient is transformed into an adjustment factor for connection strength to adjust the connection strength between content nodes and user nodes. The association weights of adjacent nodes are integrated to generate an expression vector containing association features. Extract user purpose feature value, content connection feature value, and timeliness feature value from the expression vector, construct a dynamic change weight model by combining the preference change trajectory of the historical behavior record, integrate the user purpose feature value, content connection feature value, timeliness feature value and the change weight value output by the dynamic change weight model to generate a decision vector, and perform cross-modal matching and ranking based on the decision vector to output personalized content recommendation results; The step of determining historical behavior records between different modalities in the content data stream, and generating dynamic compensation coefficients based on cross-modal transition records in the historical behavior records and the similarity correspondence table to perform deviation source analysis, includes: Read historical behavior records between different modalities of content from the content data stream in the storage system; The proportion of times text content was converted to image content in the historical behavior records was used as the text-to-image conversion value, and the proportion of times audio content was converted to video content was used as the audio-to-video conversion value. Extract the semantic relationship score and matching degree score from the similarity correspondence table, and calculate the first deviation between the text image conversion value and the semantic relationship score and the second deviation between the audio video conversion value and the matching degree score; Based on the magnitude range of the first deviation and the second deviation, a dynamic compensation coefficient positively correlated with the deviation is generated; The construction of the bipartite graph structure of content nodes and user nodes, and the conversion of the dynamic compensation coefficient into an adjustment factor for connection strength to adjust the connection strength between the content nodes and user nodes, includes: Construct a bipartite graph structure consisting of content nodes and user nodes, and assign an initial connection strength value to the connection line between each content node and the user node based on the number of historical operations of the user on the content node; The dynamic compensation coefficient is converted into an adjustment factor for the connection strength, and the initial connection strength value is calculated to obtain the adjusted connection strength value.

2. The method according to claim 1, characterized in that, The process involves extracting user purpose feature values, content relevance feature values, and timeliness feature values ​​from the expression vector, constructing a dynamic change weight model by combining the preference change trajectory of the historical behavior records, integrating the user purpose feature values, content relevance feature values, timeliness feature values, and change weight values ​​output by the dynamic change weight model to generate a decision vector, and performing cross-modal matching and ranking based on the decision vector to output personalized content recommendation results, including: The user's purpose feature value, content connection feature value, and timeliness feature value are separated from the expression vector. At the same time, the operation time series is extracted based on the preference change trajectory of historical behavior records, and the most recent operation time and operation frequency are extracted from the operation time series. A dynamic weighting model is constructed based on the recent operation time and operation frequency, and the weighting values ​​are calculated. For each content item, the user purpose feature value, the content connection feature value, and the timeliness feature value are calculated with the change weight value to generate the corresponding decision value for the content item. All decision values ​​are then organized into a decision vector according to the content item identifier corresponding to different content items. The decision values ​​of all content items in the decision vector are uniformly sorted across modalities, and the identifier of the content item with the highest ranking is output as the personalized recommendation result.

3. The method according to claim 2, characterized in that, For each content item, the user's purpose feature value, the content connection feature value, and the timeliness feature value are calculated together with the change weight value to generate a decision value for the corresponding content item. All decision values ​​are then organized into a decision vector according to the content item identifier corresponding to different content items, including: Extract the user purpose feature value, content connection feature value, and timeliness feature value corresponding to the current content item; The user's purpose feature value is multiplied by the change weight value to obtain the purpose weight value; the content connection feature value is multiplied by the change weight value to obtain the connection weight value; and the timeliness feature value is multiplied by the change weight value to obtain the timeliness weight value. The decision value for the current content item is generated by adding the target weighted value, the connection weighted value, and the timeliness weighted value. The decision value is associated with the content item identifier of the current content item, and after traversing all content items, all the decision values ​​are arranged in order of the content item identifier to form a decision vector.

4. The method according to claim 1, characterized in that, The process of integrating the association weights of neighboring nodes to generate an expression vector containing association features includes: For each user node, the adjusted connection strength value and association weight of its adjacent content nodes are obtained. The connection strength value is weighted based on the association weight to generate the association feature value of the corresponding user node. The association feature values ​​of all user nodes are arranged in node order to form an expression vector.

5. The method according to claim 1, characterized in that, The process involves analyzing the semantic relationships between text and image features, and the matching degree between sound features and video content, using a pre-established large model, to build a cross-modal similarity correspondence table. This includes: Based on the independent feature package, the text features in the text feature package and the image features in the image feature package are input into the large model. The content consistency between the text features and the image features is compared, and the semantic relationship score between each text feature and the corresponding image feature is calculated. The sound features in the sound feature package and the video content in the video content feature package are input into the large model to detect the time alignment status of the sound features and the video content, and to calculate the matching degree score between each sound feature and the corresponding video content. The semantic relationship score is recorded as a text-image relationship item according to the combination of the text feature package identifier and the image feature package identifier. The matching degree score is recorded as an audio-video relationship item according to the combination of the sound feature package identifier and the video content feature package identifier. All the text-image relationship items and the audio-video relationship items are integrated to construct a similarity correspondence table containing identifier combinations and relationship scores.

6. The method according to claim 1, characterized in that, The acquisition of a content data stream comprising multimodal feature mixtures involves configuring a multi-channel acquisition server, utilizing the hardware separation unit of the multi-channel acquisition server to decompose the content data stream into raw data streams with different modal features, and generating independent feature packages based on the feature data and identifiers corresponding to the raw data streams, including: Acquire content data streams containing text, images, audio, and video; By configuring a multi-channel acquisition server to activate multiple physical channels of the hardware separation unit, the content data stream is split to generate an original data stream containing text stream, image stream, audio stream, and video stream. Character sequences and position information are extracted from the text stream, color distribution and contour point sets are extracted from the image stream, frequency bands and loudness sequences are extracted from the sound stream, and motion trajectories and brightness changes are extracted from the video stream, so as to extract feature data corresponding to each of the original data streams; The feature data corresponding to the original data stream is appended with an identifier containing a modality type code and a timestamp, and then packaged to generate independent feature packages corresponding to text feature packages, image feature packages, sound feature packages, and video feature packages.

7. A personalized recommendation system based on a large model, characterized in that, include: The acquisition module is used to acquire a content data stream including a mixture of multimodal features. By configuring a multi-channel acquisition server, the hardware separation unit of the multi-channel acquisition server is used to decompose the content data stream into raw data streams with different modal features, and generate independent feature packages based on the feature data and identifiers corresponding to the raw data streams. The analysis module is used to analyze the semantic relationship between text features and image features, and the degree of matching between sound features and video content among multiple independent feature packages through a pre-established large model, so as to establish a cross-modal similarity correspondence table. The generation module is used to determine the historical behavior records between different modal content in the content data stream, and to generate dynamic compensation coefficients based on the cross-modal conversion records in the historical behavior records and the similarity correspondence table to perform deviation source analysis. The adjustment module is used to construct a bipartite graph structure of content nodes and user nodes, convert the dynamic compensation coefficient into an adjustment factor of connection strength to adjust the connection strength between content nodes and user nodes, and integrate the association weights of adjacent nodes to generate an expression vector containing association features. The output module is used to extract user purpose feature values, content connection feature values, and timeliness feature values ​​from the expression vector, construct a dynamic change weight model by combining the preference change trajectory of the historical behavior records, integrate the user purpose feature values, content connection feature values, timeliness feature values ​​and the change weight values ​​output by the dynamic change weight model to generate a decision vector, and perform cross-modal matching and ranking based on the decision vector to output personalized content recommendation results. The step of determining historical behavior records between different modalities in the content data stream, and generating dynamic compensation coefficients based on cross-modal transition records in the historical behavior records and the similarity correspondence table to perform deviation source analysis, includes: Read historical behavior records between different modalities of content from the content data stream in the storage system; The proportion of times text content was converted to image content in the historical behavior records was used as the text-to-image conversion value, and the proportion of times audio content was converted to video content was used as the audio-to-video conversion value. Extract the semantic relationship score and matching degree score from the similarity correspondence table, and calculate the first deviation between the text image conversion value and the semantic relationship score and the second deviation between the audio video conversion value and the matching degree score; Based on the magnitude range of the first deviation and the second deviation, a dynamic compensation coefficient positively correlated with the deviation is generated; The construction of the bipartite graph structure of content nodes and user nodes, and the conversion of the dynamic compensation coefficient into an adjustment factor for connection strength to adjust the connection strength between the content nodes and user nodes, includes: Construct a bipartite graph structure consisting of content nodes and user nodes, and assign an initial connection strength value to the connection line between each content node and the user node based on the number of historical operations of the user on the content node; The dynamic compensation coefficient is converted into an adjustment factor for the connection strength, and the initial connection strength value is calculated to obtain the adjusted connection strength value.

8. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement a personalized recommendation method based on a large model as described in any one of claims 1 to 6.

9. A computer storage medium, characterized in that, The system contains a computer program that, when executed by a computer, implements a personalized recommendation method based on a large model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Live broadcast room content identification and intelligent distribution method and system based on multi-modal fusion

    CN119377895A

  • Multi-modal user intention understanding and personalized shopping guide generation method and system

    CN120106942A