Dynamic market research sample intelligent matching method based on multi-source data fusion

By processing multi-source data through sentiment analysis, activity analysis, and purchase frequency analysis, constructing fusion vector graphs and feature matrices, and using deep neural network models for data fusion, the difficulty of multi-source data fusion is solved, and efficient survey sample matching and personalized recommendations are achieved.

CN120744520AInactive Publication Date: 2025-10-03FUZHOU AIMIFEI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510835809.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing market research methods are unable to efficiently integrate multi-source data, especially social media, player behavior data, and game purchase records, resulting in insufficient accuracy and stability in research results, and unable to meet the needs of rapid response and personalized recommendations.

Method used

Through sentiment analysis, activity analysis and purchase frequency analysis to process multi-source data, construct fusion vector graphs and fusion feature matrices, and use deep neural network models for data fusion and dynamic modeling, efficient fusion of multi-source data and dynamic behavior modeling are achieved.

Benefits of technology

It improves the accuracy of survey sample matching and the stability of recommendation results, enhances the system's responsiveness to market changes, and improves the adaptation effect of personalized recommendations and behavioral modeling tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744520A_ABST
    Figure CN120744520A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer data processing, and discloses a dynamic market research sample intelligent matching method based on multi-source data fusion, which comprises the following steps: performing sentiment analysis processing on social media data to obtain a sentiment feature vector, performing activeness analysis processing on player behavior data to obtain a behavior pattern vector, and matching the behavior pattern vector with the sentiment feature vector; the method comprises the following steps: performing purchase frequency analysis processing on a game purchase record to obtain consumption behavior vectors, uniformly mapping three types of feature vectors to the same dimension space, performing normalization and vector alignment fusion, constructing an initial fusion matrix, and further generating a fusion vector diagram, so that the relation between different data sources under a behavior tag index is subjected to structured expression, and the fusion efficiency is improved. Guiding weighted dynamic fusion processing through a fusion vector diagram, constructing a fusion feature matrix, and inputting the fusion feature matrix and the three types of original data into a deep neural network model to obtain a survey sample data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer data processing technology, and more specifically, to a dynamic market research sample intelligent matching method based on multi-source data fusion. Background Art

[0002] In existing technologies, market research often involves collecting information from multiple data sources to help understand consumer demand and market dynamics. With the popularization of channels such as social media, e-commerce platforms, and game companies, the sources of market research data have become increasingly diverse. In the animation and game industry, market research usually includes data from social media (such as player comments, forum interactions, social media posts, etc.), player behavior data (such as in-game interactions, player participation, game time, etc.), game purchase records (such as the player's history of purchasing games or in-app purchases), and player feedback (such as questionnaires, user reviews, etc.). This data not only reflects the immediate needs of players, but also has strong timeliness, personalization, and complexity of behavioral patterns. For example, players' interests and behavioral preferences may change over time, and the popularity of a certain game may suddenly increase. These factors require the market research system to be able to flexibly respond to changes and provide real-time and effective analysis.

[0003] However, existing market research methods often rely on frequent sample collection for data collection. Although this ensures data richness, it also brings problems of high cost and data redundancy. How to integrate multi-source data from different platforms without increasing the frequency of data collection and make full use of deep learning algorithms to process and analyze this data has become a difficult problem in current technology. Specifically, existing technologies find it difficult to efficiently integrate diverse data sources such as social media, player behavior data, and game purchase records, and it is difficult to improve the accuracy of dynamic market research sample matching through deep learning algorithms, thereby ensuring the stability of personalized recommendations and the efficiency of research results. Therefore, how to ensure the high reliability of research results and meet the needs of rapid response in an ever-changing market environment is a technical problem that needs to be urgently solved in the current market research field.

[0004] In view of this, the present invention proposes a dynamic market research sample intelligent matching method based on multi-source data fusion to solve the above problems. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, the present invention provides a dynamic market research sample intelligent matching method based on multi-source data fusion.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] First, it provides a dynamic market research sample intelligent matching method based on multi-source data fusion, including:

[0008] Obtain social media data, player behavior data, and game purchase records from multiple platforms, perform sentiment analysis on the social media data to obtain sentiment feature vectors, perform activity analysis on player behavior data to obtain behavior pattern vectors, and perform purchase frequency analysis on game purchase records to obtain consumption behavior vectors.

[0009] Perform multi-dimensional feature fusion processing on the emotion feature vector, behavior pattern vector, and consumption behavior vector to obtain a fusion vector graph. The fusion vector graph is used to represent the connection between social media data, player behavior data, and game purchase records in the feature space.

[0010] Perform weighted dynamic fusion processing on social media data, player behavior data, and game purchase records based on the fusion vector graph to obtain a fusion feature matrix;

[0011] The fused feature matrix, social media data, player behavior data, and game purchase records are input into the deep neural network model to obtain the survey sample dataset.

[0012] In some embodiments, a method for performing activity analysis on player behavior data to obtain a behavior pattern vector includes:

[0013] Based on the behavior records of each player in the player behavior data, the timestamps in the behavior records are sorted to obtain a behavior time series, and feature extraction processing is performed on each time point in the behavior time series to obtain the behavior feature data of the player at different time points;

[0014] Calculate the activity level of each player's behavioral feature data to obtain an activity score for each player;

[0015] The activity scores of all players are normalized to obtain a normalized activity vector, and the normalized activity vectors of all players are clustered to obtain a behavior pattern vector.

[0016] In some embodiments, a method for calculating the activity level of each player's behavioral characteristic data to obtain an activity level score for each player includes:

[0017]

[0018] Among them, H′ represents the activity score, n represents the number of behavioral feature data, and p i represents the probability distribution of the i-th behavior feature data, log(p i ) represents the logarithmic probability of the i-th behavioral characteristic data, α represents the timeliness coefficient, Δt i Represents the time difference of the i-th behavior feature data, 1+α·Δti Used to adjust the weight of behavioral feature data based on time difference.

[0019] In some embodiments, a method for performing multi-dimensional feature fusion processing on the emotion feature vector, the behavior pattern vector, and the consumption behavior vector to obtain a fusion vector graph includes:

[0020] The emotional feature vector, behavioral pattern vector, and consumer behavior vector are mapped to the same dimensional space and normalized to obtain three types of normalized feature vectors with consistent structures. Each dimensional element in each normalized feature vector corresponds to a preset behavioral label index.

[0021] Performing vector alignment fusion processing on the three types of normalized eigenvectors to construct an initial fusion matrix, wherein the vector alignment fusion processing refers to performing weighted splicing processing on the three types of eigenvectors in each dimension according to the index order;

[0022] The initial fusion matrix is ​​reconstructed by feature density to obtain a fusion density matrix. The feature dimension corresponding to each behavior label index in the fusion density matrix is ​​used as a node, and the joint change intensity of any two feature dimensions in the initial fusion matrix is ​​used as an edge to construct a fusion vector graph.

[0023] In some embodiments, a method of performing feature density reconstruction processing on an initial fusion matrix to obtain a fusion density matrix includes:

[0024] Based on the feature dimension element value corresponding to each behavior label index in the initial fusion matrix, calculate its joint occurrence frequency in the three types of normalized feature vectors;

[0025] All joint occurrence frequencies are organized into a vector structure according to the index order to obtain the original density vector, which is used to represent the joint response frequency of each feature dimension;

[0026] Perform time series window sliding processing on the original density vector, evolve the sequence window according to the preset behavior label, perform local mean and range analysis on the density values ​​within each window range, and generate a set of density fluctuation factors;

[0027] The original density vector is nonlinearly smoothed according to the density fluctuation factor set to generate a fused density matrix.

[0028] In some embodiments, a method for performing local mean and range analysis on density values ​​within each window range to generate a set of density fluctuation factors includes:

[0029] The density value sequence of each feature dimension in the fusion density matrix is ​​recorded as Q = {q1,q2,...,q L}, based on the window length u and step size v, Q is divided into sliding windows to obtain the density subsequence set R = {Q j}, where each Q j ={q j ,q j+1 ,...,q j+u-1}, satisfying j∈{1,1+v,...,L-u+1};

[0030] For each Q j , calculate its second-order disturbance, based on the second-order disturbance, and introduce the exponential adjustment factor, construct the density fluctuation factor, and form all density fluctuation factors into a density fluctuation factor set.

[0031] In some embodiments, for each Q j , the methods for calculating its second-order perturbation include:

[0032]

[0033] Among them, j represents the second-order perturbation corresponding to the jth window, u represents the length of the sliding window, u-1 represents the number of time points in each window that can be used to calculate the perturbation, and q j+k represents the density value at the j+kth time point, q j+k-1 Indicates the density value at the previous time point, (q j+k -q j+k-1 ) 2 represents the square of the difference between two moments, Indicates averaging of the sum of squared disturbances.

[0034] In some embodiments, a method for constructing a density fluctuation factor based on a second-order perturbation and introducing an exponential adjustment factor includes:

[0035]

[0036] Among them, θ j represents the density fluctuation factor corresponding to the j-th window, represents the cumulative result of all squares of density values ​​in the jth window, ∈ represents the smoothing constant, exp(-λ·σ j ) represents the standard deviation of the disturbance σ j The exponential adjustment factor constructed, where λ represents the adjustment coefficient, σ j represents the standard deviation of the density value in the j-th window,

[0037] In some embodiments, a method for performing weighted dynamic fusion processing on social media data, player behavior data, and game purchase records based on a fusion vector graph to obtain a fusion feature matrix includes:

[0038] Obtain the node set and edge set in the fusion vector graph, where the node set represents the feature dimension corresponding to each line label index, and the edge set represents the joint change strength between any two feature dimensions;

[0039] Based on the behavioral label index order of each type of original feature vector in social media data, player behavior data, and game purchase records, the three types of original feature vectors are mapped to the node set of the fusion vector graph to generate three types of index-bound feature vectors;

[0040] Normalize the weights of the edge sets in the fusion vector graph to obtain the normalized edge weight vector. Perform dot product processing on the normalized edge weight vector and the corresponding elements in the three types of index binding feature vectors to generate the fusion response vector.

[0041] Perform interval segmentation processing on the fusion response vector, map the corresponding index to the fusion level label according to the interval to which the response value belongs, and combine all indexes in the level label order to generate a level label set;

[0042] A fusion feature matrix is ​​constructed based on the level label set and the fusion response value corresponding to each index in the fusion response vector.

[0043] In some embodiments, a method for generating a fused response vector by performing dot product processing on the normalized edge weight vector and corresponding elements in the three-category index binding feature vector includes:

[0044] According to the frequency of occurrence of the feature dimension connected by each edge in the normalized initial vector in the three types of index-bound feature vectors, a set of dynamic adjustment factors is constructed, where each dynamic adjustment factor is used to reflect the local sensitivity of the feature dimension to the change of the response value;

[0045] Perform element-wise product processing on the normalized edge weight initial vector and the adjustment factor at the corresponding position in the dynamic adjustment factor set to generate a fused edge weight vector;

[0046] The fusion response vector is constructed by performing weighted summation processing on the fusion edge weight vector and the values ​​of the corresponding feature dimensions in the three-category index binding feature vector;

[0047] According to the order of weighted values ​​corresponding to all behavior label indexes in the fusion response vector, the column structure of the fusion feature matrix is ​​organized and generated to obtain the fusion feature matrix.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] The present invention realizes efficient fusion and dynamic modeling of social media data, player behavior data and game purchase records by constructing a fusion vector graph and a fusion feature matrix, thereby effectively solving the problems of data fusion difficulties and insufficient research sample matching accuracy caused by diverse data sources and limited collection frequency in the existing technology. By performing sentiment analysis on social media data, sentiment feature vectors are obtained, activity analysis is performed on player behavior data to obtain behavior pattern vectors, and purchase frequency analysis is performed on game purchase records to obtain consumption behavior vectors. The three types of feature vectors are uniformly mapped to the same dimensional space, normalized and vector-aligned fused, and an initial fusion matrix is ​​constructed, thereby generating a fusion vector graph, so that the connection between different data sources under the behavior label index is structured. Then, the weighted dynamic fusion processing is guided by the fusion vector graph to construct a fusion feature matrix. On this basis, the fusion feature matrix and the original three types of data are input into the deep neural network model together to obtain the survey sample data set. This realizes the structured expression and dynamic behavior modeling of multi-source data without increasing the frequency of data collection. This scheme improves the accuracy of survey sample matching and the stability of recommendation results, enhances the system's responsiveness to market changes, and significantly improves the adaptation effect of the survey system in personalized recommendation and behavior modeling tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 Schematic diagram of the process of the dynamic market research sample intelligent matching method based on multi-source data fusion in the present invention;

[0051] Figure 2 The figure is a schematic structural diagram of an electronic device in the present invention. DETAILED DESCRIPTION

[0052] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings. In the following detailed description, many specific details are set forth to provide a thorough understanding of the described exemplary embodiments. However, it is obvious to those skilled in the art that the described embodiments can be practiced without some or all of these specific details. In other exemplary embodiments, well-known structures are not described in detail to avoid unnecessarily obscuring the concepts of the present disclosure. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. At the same time, the various aspects described in the embodiments can be arbitrarily combined without conflict.

[0053] Example 1

[0054] See also Figure 1 As shown, this embodiment discloses a method for intelligently matching dynamic market research samples based on multi-source data fusion, including:

[0055] S10: Acquire social media data, player behavior data, and game purchase records from multiple platforms, perform sentiment analysis on the social media data to obtain sentiment feature vectors, perform activity analysis on the player behavior data to obtain behavior pattern vectors, and perform purchase frequency analysis on the game purchase records to obtain consumption behavior vectors;

[0056] In this embodiment, the multiple platforms refer to social media platforms, game behavior monitoring platforms, and e-commerce platforms, which respectively provide players' comment interactions, game participation, and purchase history data. Social media platforms refer to mainstream Chinese social platforms including but not limited to Weibo, WeChat, and Tik Tok, which provide comments, posts, interaction data, and other content posted by players to reflect players' emotional attitudes, social behaviors, and real-time preferences.

[0057] In-game behavior monitoring platforms refer to gaming platforms and behavior monitoring systems in China, such as Tencent's WeGame platform and NetEase's NetEase Cloud Gaming platform, which collect player behavior data in the game, including game duration, task completion status, interactive behavior, etc., to reflect player activity, participation level, and gaming preferences. E-commerce platforms refer to Chinese e-commerce platforms including but not limited to Taobao, JD.com, Pinduoduo, etc., which provide players' purchase records and in-app purchase history data to reflect players' consumption habits, purchase frequency, consumption amount and other information.

[0058] Methods for performing sentiment analysis on social media data and obtaining sentiment feature vectors include:

[0059] Based on each comment text in the social media data, the comment text is segmented to obtain the segmentation results, and each segmentation result is filtered for stop words to obtain a cleaned segmentation set;

[0060] Perform sentiment polarity analysis on the cleaned word set to obtain the sentiment score of each word, and remove words with sentiment scores lower than the preset score threshold to obtain valid word segments;

[0061] The effective word segmentations in each comment text are weighted and summed to obtain the sentiment score vector of the comment text. The sentiment score vectors of all comment texts are statistically analyzed to obtain the sentiment feature vector.

[0062] In this embodiment, in the first step, word segmentation is performed on the comment text in social media data. The principle is to convert natural language text into a structured set of terms that can be used for subsequent analysis. Usually, Chinese word segmentation algorithms based on dictionary matching or probability models are used, such as maximum forward matching, hidden Markov model, etc. The purpose of word segmentation is to provide a finer input unit for subsequent stop word filtering and sentiment polarity scoring, ensuring that the granularity of sentiment judgment reaches the word level rather than the sentence level. After obtaining the preliminary word segmentation result, stop word filtering needs to be performed. This processing is based on a predefined stop word list. Function words such as "de", "shi", "he", etc. lack information in sentiment judgment and are therefore excluded. The cleaned word segmentation set after exclusion can focus more on words with emotional colors and provide effective input for sentiment polarity analysis, avoiding interference from invalid terms on the scoring result.

[0063] It should be noted that performing sentiment polarity analysis on the cleaned word segmentation set is the core link of this step. The principle is to quantitatively score according to the emotional intensity value of each word in the sentiment dictionary. Typical dictionaries include NTUSD, HowNet, etc. Each word is assigned a positive or negative score. The purpose of this processing is to establish a numerical mapping of sentiment semantics, enabling the text to be converted into a vector representation to support subsequent weighting and statistical processing. Words with sentiment scores lower than the preset threshold are excluded. The principle of this step is to introduce a noise control mechanism and only retain words with significant emotional intensity to participate in vector calculation. For example, if the threshold is set to 0.2 and the sentiment score of a certain word is 0.1, it will be excluded. The purpose of the processing is to improve the representativeness and stability of the sentiment feature vector and reduce the perturbation of marginal words on the statistical value.

[0064] For example, in a comment "This game has exquisite graphics but sluggish operation", after word segmentation, we get "game", "graphics", "exquisite", "operation", "sluggish". Among them, "exquisite" may obtain a positive score of +0.8, and "sluggish" obtains a negative score of -0.7. After weighted summation of these two words, the sentiment score vector of this comment is [+0.8, -0.7]. Subsequently, statistical processing is performed on all comment sentiment vectors, such as extraction of mean value, variance, maximum and minimum values, to construct a sentiment feature vector with a unified dimension to describe the overall trend of emotions in social media data.

[0065] The methods for performing activity analysis on player behavior data to obtain the behavior pattern vector include:

[0066] Based on the behavior records of each player in the player behavior data, sort the timestamps in the behavior records to obtain a behavior time series, and perform feature extraction processing on each time point in the behavior time series to obtain the behavior feature data of the player at different time points;

[0067] Calculate the activity level of each player's behavioral feature data to obtain an activity score for each player;

[0068] The activity scores of all players are normalized to obtain a normalized activity vector, and the normalized activity vectors of all players are clustered to obtain a behavior pattern vector.

[0069] First, based on the behavior records of each player in the player behavior data, the timestamps in the behavior records are sorted. The principle is to construct a behavior time series so that the behavior data has time continuity and an analysis basis. The purpose of this step is to provide time series structure support for subsequent feature extraction to ensure that each behavior has a clear time location. It can be understood that after completing the time series construction, feature extraction processing is performed on each time point in order to extract the player's behavior performance within the time segment. The behavioral feature data usually extracted include login status, operation type, number of tasks completed, and interaction frequency, and are bound to the behavior label and timestamp to form a multi-dimensional behavior feature vector. The purpose of this step is to structure the original behavior record for subsequent calculation and analysis.

[0070] It should be noted that the activity calculation and processing of behavioral feature data is a key link in the processing process. The principle is to perform weighted accumulation of the number of player behaviors, behavior types, and behavior duration within a specified time window to obtain an activity score used to measure the level of behavioral activity. The purpose of this step is to establish a numerical activity indicator for each player and standardize the activity scores of all players. Standardization refers to converting the characteristic value into a distribution with a mean of 0 and a standard deviation of 1. The Z-score method is often used. The purpose of this step is to reduce the impact of differences in the magnitude of player behavior on subsequent analysis and improve the consistency and comparability of subsequent processing.

[0071] For example, during a certain game cycle, player A logs in twice a day and performs an average of five interactive behaviors each time, while player B logs in once a day but performs an average of 10 task operations. The former reflects stable and frequent participation, while the latter reflects high-density concentrated operations. In the activity score, they can be expressed as a weighted combination of behavior persistence and intensity, which can be formed into an activity vector under a unified scale after standardization.

[0072] Finally, the standardized activity vectors of all players are clustered. Common clustering methods include K-means, DBSCAN, or hierarchical clustering. The purpose of this clustering is to classify players with similar behavior patterns into the same category. Representative behavior pattern vectors are formed through cluster centers or intra-category statistical features for subsequent behavior prediction, user profile construction, or personalized recommendation model input.

[0073] The method for calculating the activity level of each player's behavioral characteristic data to obtain the activity level score of each player includes:

[0074]

[0075] Among them, H′ represents the activity score, n represents the number of behavioral feature data, and p i represents the probability distribution of the i-th behavior feature data, log(p i ) represents the logarithmic probability of the i-th behavioral characteristic data, α represents the timeliness coefficient, Δt i Represents the time difference of the i-th behavior feature data, 1+α·Δt i Used to adjust the weight of behavioral feature data based on time difference.

[0076] Methods for analyzing the purchase frequency of game purchase records to obtain consumption behavior vectors include:

[0077] Sort each purchase record data in the game purchase record by time to obtain a purchase time series; calculate the purchase interval of each purchase record data in the purchase time series to obtain the time interval data of each purchase;

[0078] Perform frequency analysis on the data of each time interval to obtain a purchase frequency score, perform normalization on the purchase frequency scores of all players to obtain a normalized consumption behavior vector, and perform clustering on all consumption behavior vectors to obtain a consumption behavior vector.

[0079] It should be noted that α is generally set to 0.1-0.3 to control the time decay rate. Frequency analysis of each time interval data is the key processing link of this step. The principle is to calculate the number of purchases per unit time or analyze the trend of changes in purchase intervals based on statistical laws or sliding time windows, and construct a frequency score indicator. For example, if a player's purchase time is the 1st, 4th, 10th and 11th day, then the purchase time intervals are 3 days, 6 days and 1 day. A statistical window of 7 days can be set, and the number of purchases in the window is counted as 2 times, which can be expressed as a window frequency score of 2. If weighted average processing is used, a time decay factor can be further introduced to give a higher weight to the most recent purchase behavior, thereby calculating the player's final purchase frequency score, which serves as the core quantitative basis for measuring the player's consumption activity.

[0080] S20: Perform multi-dimensional feature fusion processing on the emotion feature vector, behavior pattern vector, and consumption behavior vector to obtain a fusion vector graph. The fusion vector graph is used to represent the connection between social media data, player behavior data, and game purchase records in the feature space;

[0081] The method of performing multi-dimensional feature fusion processing on the emotion feature vector, the behavior pattern vector, and the consumption behavior vector to obtain a fusion vector graph includes:

[0082] The emotional feature vector, behavioral pattern vector, and consumer behavior vector are mapped to the same dimensional space and normalized to obtain three types of normalized feature vectors with consistent structures. Each dimensional element in each normalized feature vector corresponds to a preset behavioral label index.

[0083] Performing vector alignment fusion processing on the three types of normalized eigenvectors to construct an initial fusion matrix, wherein the vector alignment fusion processing refers to performing weighted splicing processing on the three types of eigenvectors in each dimension according to the index order;

[0084] The initial fusion matrix is ​​reconstructed by feature density to obtain a fusion density matrix. The feature dimension corresponding to each behavior label index in the fusion density matrix is ​​used as a node, and the joint change intensity of any two feature dimensions in the initial fusion matrix is ​​used as an edge to construct a fusion vector graph.

[0085] In this embodiment, the emotional feature vector, behavioral pattern vector and consumption behavior vector are first mapped to the same dimensional space respectively. The principle is to make the three types of features have the same dimensional structure and index alignment relationship through vector transformation operation, so as to ensure that the subsequent fusion process has a consistency basis. The mapping operation usually relies on a preset behavioral label index structure, which serves as a unified reference in the three types of features. The purpose is to ensure that each dimension clearly points to a specific behavioral meaning. For example, index 0 can represent "high-frequency login", index 1 represents "positive evaluation", and index 2 represents "high-value purchase". After completing the dimensional mapping, the three types of feature vectors are normalized to eliminate the differences in numerical distribution and dimensional units of features from different sources. Commonly used methods are Z-score normalization or Min-Max normalization to ensure that the element values ​​on each dimension can be fairly weighted. The normalization result forms a feature input sequence with consistent structure, laying a numerical foundation for subsequent alignment fusion.

[0086] It should be noted that the vector alignment fusion processing based on the three types of normalized feature vectors is a key step in constructing the fusion information structure. The principle is to weight the element values ​​of the three feature vectors based on the same index position according to the preset weights to form a fusion vector. For example, if the normalized values ​​of a certain index dimension in the three types of features are 0.6, 0.4 and 0.8 respectively, and the fusion weights are set to 0.3, 0.5 and 0.2 respectively, the corresponding fusion value is calculated as: 0.6 multiplied by 0.3, plus 0.4 multiplied by 0.5, plus 0.8 multiplied by 0.2, and the final result is 0.52. Similarly, the fusion processing is completed for all behavior label index dimensions to obtain an initial fusion matrix with a consistent structure.

[0087] Furthermore, the feature density reconstruction processing of the initial fusion matrix is ​​performed in order to identify highly coupled behavioral structures through the change correlation between adjacent dimensions. The principle is to calculate the degree of coordinated change between any two dimensions under multiple player samples. For example, if the numerical fluctuations of index 3 and index 7 in multiple samples are similar, then their joint change intensity is relatively high, and they can be set as an edge with the edge weight being the change correlation value. In this way, a fusion vector graph is constructed, in which each feature dimension is a node and the joint change intensity is an edge. The overall output fusion vector graph structure has a behavior label cascade relationship.

[0088] Methods for performing feature density reconstruction on the initial fusion matrix to obtain a fusion density matrix include:

[0089] Based on the feature dimension element value corresponding to each behavior label index in the initial fusion matrix, calculate its joint occurrence frequency in the three types of normalized feature vectors;

[0090] All joint occurrence frequencies are organized into a vector structure according to the index order to obtain the original density vector, which is used to represent the joint response frequency of each feature dimension;

[0091] Perform time series window sliding processing on the original density vector, evolve the sequence window according to the preset behavior label, perform local mean and range analysis on the density values ​​within each window range, and generate a set of density fluctuation factors;

[0092] The original density vector is nonlinearly smoothed according to the density fluctuation factor set to generate a fused density matrix.

[0093] Methods for performing local mean and range analysis on density values ​​within each window range and generating a set of density fluctuation factors include:

[0094] The density value sequence of each feature dimension in the fusion density matrix is ​​recorded as Q = {q1,q2,...,q L}, based on the window length u and step size v, Q is divided into sliding windows to obtain the density subsequence set R = {Q j}, where each Q j ={q j ,q j+1 ,...,q j+u-1}, satisfying j∈{1,1+v,...,L-u+1};

[0095] For each Q j , calculate its second-order disturbance, based on the second-order disturbance, and introduce the exponential adjustment factor, construct the density fluctuation factor, and form all density fluctuation factors into a density fluctuation factor set.

[0096] For each Q j, the methods for calculating its second-order perturbation include:

[0097]

[0098] Among them, j represents the second-order perturbation corresponding to the jth window, u represents the length of the sliding window, u-1 represents the number of time points in each window that can be used to calculate the perturbation, and q j+k represents the density value at the j+kth time point, q j+k-1 Indicates the density value at the previous time point, (q k+k -q j+k-1 ) 2 represents the square of the difference between two moments, Indicates averaging of the sum of squared disturbances.

[0099] Based on the second-order perturbation and introducing the exponential adjustment factor, the method of constructing the density fluctuation factor includes:

[0100]

[0101] Among them, θ j represents the density fluctuation factor corresponding to the j-th window, represents the cumulative result of all squares of density values ​​in the jth window, ∈ represents the smoothing constant, exp(-λ·σ j ) represents the standard deviation of the disturbance σ j The exponential adjustment factor constructed, where λ represents the adjustment coefficient, σ j represents the standard deviation of the density value in the jth window,

[0102] In this embodiment, in order to further explain clearly, it is first necessary to calculate the joint occurrence frequency of the feature dimension element value corresponding to each behavior label index in the initial fusion matrix in the three types of normalized feature vectors. The core purpose is to count the response activity and interconnection coefficient of each multi-source feature combination behavior calibration. For example, for the elements with values ​​q1, q2, and q3 at index j, if their values ​​fall in any discrete interval with a number of values ​​of 4, such as [0,0.5), [0.5,1), [1,1.5), [1.5,2), then the joint occurrence frequency of the index on the sample is 4, and the three values ​​fall into the 1st, 2nd, and 3rd intervals respectively, and the corresponding joint discrete coordinates are (1,2,3), then the joint occurrence frequency on the index can be recorded as 1; if the same discrete combination appears repeatedly on the same index in multiple subsequent samples, the frequency is accumulated accordingly.

[0103] After the joint occurrence frequency statistics are completed, the frequency values ​​corresponding to all indexes need to be organized into a vector structure according to the order of the behavior label index to obtain the original density vector. The principle is: through the joint response statistics of all feature dimensions, a linearly arranged density sequence is constructed to describe the response intensity distribution of each dimension; for example, if there are 5 behavior label indexes in total, and the joint occurrence frequencies obtained by statistics are 4, 3, 2, 5, and 1, respectively, then the original density vector can be recorded as D = {4, 3, 2, 5, 1}. The larger the value of each element in the vector, the higher the co-occurrence density of the dimension in multiple samples, reflecting that the feature dimension has a stronger response consistency in multi-source fusion. The dynamic perturbation feature calculation will be performed based on the original density vector in the future.

[0104] The density values ​​within the window range are subjected to local mean and range analysis to generate a set of density fluctuation factors. The principle of this processing is: by setting a sliding window of fixed length, a local subsequence in the original density vector is extracted, and the local disturbance level and distribution structure of the subsequence median are used to construct a disturbance feature, which is used to measure the dynamic change intensity of density between different behavior segments. For example, if the original density vector is D = {4, 3, 2, 5, 1}, the window length u = 3, and the step size v = 1, then three density subsequences Q1 = {4, 3, 2}, Q2 = {3, 2, 5}, and Q3 = {2, 5, 1} can be generated. Each subsequence will subsequently be subjected to disturbance analysis and calculation to obtain density fluctuation factors θ1, θ2, θ3, etc.

[0105] Specifically, after completing the density subsequences Q1 = {4, 3, 2}, Q2 = {3, 2, 5}, and Q3 = {2, 5, 1}, the corresponding perturbations ζ1, ζ2, and ζ3 are calculated based on the perturbation formula, and the parameters λ = 0.3 and the standard deviation σ are introduced. j Based on the smoothing term ε=0.1, the density fluctuation factors of each window are calculated using the above mathematical model, which are θ1=0.210, θ2=0.178, and θ3=0.154 respectively.

[0106] For example, assuming the total number of behavior labels is 5 and the original density vector is D = {4, 3, 2, 5, 1}, the fused density matrix M is a 5 × 5 symmetric matrix, where the elements on the main diagonal position directly use the density value of the corresponding index in the original density vector, and the elements on the off-diagonal position are based on whether the corresponding two behavior labels are both included in a sliding window. If so, the weighted average density value is filled in and adjusted in combination with the density fluctuation factor within the window. If the two label indexes are not covered by the same window, the corresponding position element is filled with 0, indicating that there is no direct coupling relationship. Therefore, the fused density matrix M is expressed as follows:

[0107]

[0108] First, based on a fixed-length sliding window, the density value sequence on each feature dimension is divided into several continuous local subsequences. Each subsequence represents the local density structure at adjacent time points or sample dimensions. Then, the perturbation amplitude is calculated for each subsequence, and the rate of change between adjacent density values ​​is extracted. Further, combined with the overall fluctuation level within the subsequence, a disturbance index reflecting the activeness of density changes is comprehensively constructed. This index takes into account both the jump intensity between density values ​​and the distribution breadth of the jump within the window. To avoid the interference effect caused by a single high or low density value, this method introduces a normalization mechanism based on the overall density intensity. In addition, a dynamic adjustment factor is introduced based on the fluctuation distribution characteristics within the window. This adjustment factor has nonlinear suppression capabilities, which can enhance the response in high-fluctuation areas and suppress the response in low-fluctuation areas. Therefore, the final fluctuation factor result can better reflect the dynamics and heterogeneity of behavioral responses.

[0109] Compared with the existing technology that only makes judgments based on single-point intensity or linear statistical features, the innovation of this method lies in: by introducing the disturbance change rate and exponential adjustment mechanism, the calculation of the fluctuation factor is made more dynamically adaptable and responsive, thereby improving the robustness and discrimination ability of the fusion density structure when processing cross-source and cross-dimensional behavioral data, and significantly enhancing the edge weight recognition accuracy between behavioral labels in the subsequent fusion vector graph and the definition of behavioral sensitive segments.

[0110] S30: Perform weighted dynamic fusion processing on the social media data, player behavior data, and game purchase records according to the fusion vector graph to obtain a fusion feature matrix;

[0111] Methods for performing weighted dynamic fusion processing on social media data, player behavior data, and game purchase records based on the fusion vector graph to obtain a fusion feature matrix include:

[0112] Obtain the node set and edge set in the fusion vector graph, where the node set represents the feature dimension corresponding to each line label index, and the edge set represents the joint change strength between any two feature dimensions;

[0113] Based on the behavioral label index order of each type of original feature vector in social media data, player behavior data, and game purchase records, the three types of original feature vectors are mapped to the node set of the fusion vector graph to generate three types of index-bound feature vectors;

[0114] Normalize the weights of the edge sets in the fusion vector graph to obtain the normalized edge weight vector. Perform dot product processing on the normalized edge weight vector and the corresponding elements in the three types of index binding feature vectors to generate the fusion response vector.

[0115] Perform interval segmentation processing on the fusion response vector, map the corresponding index to the fusion level label according to the interval to which the response value belongs, and combine all indexes in the level label order to generate a level label set;

[0116] A fusion feature matrix is ​​constructed based on the level label set and the fusion response value corresponding to each index in the fusion response vector.

[0117] The method of performing dot product processing on the normalized edge weight vector and the corresponding elements in the three-category index binding feature vector to generate a fused response vector includes:

[0118] According to the frequency of occurrence of the feature dimension connected by each edge in the normalized initial vector in the three types of index-bound feature vectors, a set of dynamic adjustment factors is constructed, where each dynamic adjustment factor is used to reflect the local sensitivity of the feature dimension to the change of the response value;

[0119] Perform element-wise product processing on the normalized edge weight initial vector and the adjustment factor at the corresponding position in the dynamic adjustment factor set to generate a fused edge weight vector;

[0120] The fusion response vector is constructed by performing weighted summation processing on the fusion edge weight vector and the values ​​of the corresponding feature dimensions in the three-category index binding feature vector;

[0121] According to the order of weighted values ​​corresponding to all behavior label indexes in the fusion response vector, the column structure of the fusion feature matrix is ​​organized and generated to obtain the fusion feature matrix.

[0122] In this embodiment, in order to construct a fusion vector graph and realize the fusion of three types of features, it is necessary to first obtain a set of behavior label indexes and their corresponding structural connection relationships. The behavior label index set is recorded as {S1, S2, S3, S4, S5}, which respectively correspond to the five dimensions in the original density vector D = {4, 3, 2, 5, 1}, where each index S j To represent a fusion dimension from social media data, player behavior data, or game purchase records, and to further characterize the joint change relationship between the dimensions, it is necessary to combine the sliding window sequence of the original density vector as follows:

[0123] Q1 = {A1, A2, A3}, Q2 = {A2, A3, A4} and Q3 = {A3, A4, A5}, construct an edge set in the graph, where undirected edges are established between any two index nodes in each group of windows. If the corresponding window perturbation factors are θ1 = 0.210, θ2 = 0.178, and θ3 = 0.154, a graph structure containing six edges can be formed, and each edge is assigned a perturbation factor as the initial edge weight to represent the intensity of the local density response.

[0124] It can be understood that after the graph structure is established, in order to enable each node in the fusion graph to receive input information from the three types of feature sources, the values ​​of the corresponding index positions in the social media feature vector, player behavior feature vector and game purchase feature vector need to be mapped to nodes A1 to A5 in the node set in sequence. For example, the emotional feature vector value is {0.6, 0.3, 0.8, 0.2, 0.1}, the behavioral feature vector value is {0.5, 0.4, 0.9, 0.3, 0.2}, and the consumption feature vector value is {0.7, 0.2, 0.6, 0.1, 0.3}. Then the A1 node is bound to three types of values ​​0.6, 0.5 and 0.7 respectively, the A2 node is bound to 0.3, 0.4 and 0.2, and so on, so as to ensure that each graph node has three types of information input.

[0125] After binding the three types of feature values ​​within the node set, the fusion weights for each edge must be calculated based on the edge set within the graph structure. The principle behind this process is to weight the three types of feature values ​​bound to the two nodes connected by the edge, and then use the edge's perturbation factor as an overall adjustment coefficient to express the edge's structural strength in the three-type information fusion process. For example, the edge (A1, A2) connecting the nodes within the Q1 window has a perturbation factor of θ1 = 0.210, with A1 binding values ​​of 0.6, 0.5, and 0.7, and A2 binding value of 0. 3, 0.4, 0.2. Assuming that the fusion weights of the three types of features are w1=0.4, w2=0.4, and w3=0.2 respectively, the weight of the edge in the fusion edge weight vector is calculated as θ1×[(0.6+0.3)×0.4+(0.5+0.4)×0.4+(0.7+0.2)×0.2]=0.210×[(0.9)×0.4+(0.9)×0.4+(0.9)×0.2]=0.210×0.9=0.189, indicating that the fusion strength of the edge under the joint influence of the three types of information is 0.189.

[0126] It can be understood that after the construction of the fusion edge weight vector is completed, a weighted summation process is performed based on the fusion edge weight vector and the value of the corresponding feature dimension in the three-category index binding feature vector, which is used to further project the fusion influence of each edge in the graph structure to each node and construct the final fusion response vector. The principle is: the fusion edge weights of all edges connected to each node are weightedly superimposed according to the eigenvalues ​​in the connection direction, so as to measure the centralized response level of the node in the entire fusion graph, reflecting the comprehensive significance of multi-source features on the label index.

[0127] Furthermore, the column structure of the fusion feature matrix is ​​organized and generated according to the order of weighted values ​​corresponding to all behavioral label indexes in the fusion response vector in order to convert the one-dimensional fusion results into a two-dimensional structured expression. This process sorts and groups all indexes according to the fusion response values, divides them into a preset level label set, and organizes them into the column structure of the fusion feature matrix in order of level. This feature matrix serves as the basis for subsequent model input and has the ability to summarize and express the fusion relationship of multiple sources of behavior. It is suitable for application scenarios such as survey sample data set generation or personalized recommendation modeling.

[0128] It is understandable that the weighted dynamic fusion process driven by the fusion vector graph, after completing the construction of the fusion feature matrix, can significantly improve the collaborative expression ability of multi-source heterogeneous features in behavior recognition tasks, which is specifically reflected in the following three aspects:

[0129] First, the fusion process extracts the joint perturbation relationship between behavior label indexes through a sliding window method, and establishes a structural graph model with local density response perception capabilities, thereby effectively identifying areas sensitive to changes in the original density distribution. This structured modeling method can capture the potential collaborative change relationship between labels that is difficult to identify in traditional methods, and improve the accuracy of correlation characterization between behavior labels.

[0130] Secondly, the edge weight construction process in the fusion vector graph introduces a joint adjustment mechanism for three types of feature sources. On the basis of maintaining the independence of their respective feature structures, it completes the aggregate modeling of multi-source inputs under the same label index. This process does not rely on high-frequency sampling or redundant feature stacking, but instead performs weighted fusion based on the structural connection relationship between behavioral label indexes, ensuring a balance between dimensionality control and expressiveness of the fused feature matrix.

[0131] Finally, the mapping process of the fused response vector to the fused feature matrix is ​​achieved through the division of hierarchical labels and the sorting of column structures, so that the final generated fused feature matrix has obvious behavioral discrimination capabilities and feature difference expression effects, which can improve the model's discrimination performance and sample adaptability in subsequent tasks. It is particularly suitable for multi-source data mining scenarios that require high-dimensional feature compression, fine modeling of behavioral labels, or enhanced structural perception.

[0132] S40: Input the fused feature matrix, social media data, player behavior data, and game purchase records into the deep neural network model to obtain the survey sample dataset.

[0133] It can be understood that a survey sample dataset refers to a type of structured data set with labeled or implicit preference relationships constructed to support tasks such as user preference modeling, behavior prediction, or personalized recommendation. This dataset usually contains multi-source behavior information and associated labels of target users or groups, which are used to describe the user's behavioral feature expression and decision-making tendencies in specific task scenarios, such as "the strength of interest in a certain type of game" and "the willingness to click on specific content". Unlike ordinary behavioral raw data, the survey sample dataset has clear structure, compressed features, and discriminative capabilities. It is the result of multi-source data after structural transformation and behavior-driven modeling, and serves the efficient training and generalization evaluation of downstream models.

[0134] Furthermore, in this embodiment, the fused feature matrix is ​​a structural expression result constructed by guiding the construction of original features such as social media data, player behavior data, and game purchase records through a fusion vector graph. Its core value lies in eliminating the dimensional differences and redundant information interference between the original multi-source data, and enhancing the correlation expression between features through edge weight adjustment and response sorting mechanisms. Therefore, when the fused feature matrix, social media data, player behavior data, and game purchase records are jointly fed into the deep neural network model as input, the model can automatically learn the multidimensional joint expression pattern of the three types of data under the behavioral label index.

[0135] Specifically, the deep neural network model can include the following three layers:

[0136] Structured feature layer: The fused feature matrix is ​​used as the first layer input, providing a standardized and correlation-enhanced feature expression. Each row of the matrix corresponds to a user sample, and each column corresponds to the fused behavior label feature. It can be directly input into the fully connected layer for high-order feature extraction.

[0137] Detailed feature layer: The raw data (social media data, player behavior data, and game purchase records) are processed by independent embedding layers or encoders (such as LSTM and CNN) and converted into fixed-dimensional feature vectors, retaining the fine-grained information of the original data.

[0138] Feature concatenation layer: In the first hidden layer of the model, the output of the fused feature matrix is ​​concatenated with the three types of original feature encoding vectors to form a composite feature representation that combines global structure and local details.

[0139] The principle is as follows: the fused feature matrix provides the structural basis after the initial fusion, while the original three types of feature data serve as input to supplement the model's sensitivity to individual behavioral details; the deep neural network extracts and transforms these four types of inputs layer by layer through a nonlinear mapping structure, so that the network gradually converges to a set of high-order feature representations that can characterize individual behavioral tendencies and have the ability to distinguish during the training process. Finally, these high-order feature outputs form the sample items in the survey sample data set. Each item not only retains the expressive richness of multi-source data, but also has structural consistency under the fusion response dimension, and can be directly used as input for classification, prediction, sorting and other models. The construction process of the survey sample data set is essentially to compress the original multi-source features into a highly expressive feature matrix through a fusion graph structure, and then the deep network completes the behavioral semantic mapping and expression reconstruction process. It is a data generation method driven by the collaborative fusion of perception, structural reinforcement and representation learning. This method can significantly improve downstream modeling efficiency, enhance sample quality, and have good cross-platform adaptability.

[0140] Example 2

[0141] like Figure 2 As shown, this embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the dynamic market research sample intelligent matching method based on multi-source data fusion provided by the above methods.

[0142] The detailed description set forth above in conjunction with the accompanying drawings describes examples and does not represent all examples that can be implemented or fall within the scope of the claims. The terms "example" and "exemplary" when used in this specification mean "used as an example, instance or illustration" and do not mean "better than or better than other examples."

[0143] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Therefore, use of these phrases may refer to more than just one embodiment, and further, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0144] It should also be noted that these embodiments may be described as a process depicted as a flowchart, structure diagram, or block diagram, and that although the flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently, and the order of the operations may be rearranged.

Claims

1. A dynamic market research sample intelligent matching method based on multi-source data fusion, characterized by: include: Obtain social media data, player behavior data, and game purchase records from multiple platforms, perform sentiment analysis on the social media data to obtain sentiment feature vectors, perform activity analysis on player behavior data to obtain behavior pattern vectors, and perform purchase frequency analysis on game purchase records to obtain consumption behavior vectors. Perform multi-dimensional feature fusion processing on the emotion feature vector, behavior pattern vector, and consumption behavior vector to obtain a fusion vector graph. The fusion vector graph is used to represent the connection between social media data, player behavior data, and game purchase records in the feature space. Perform weighted dynamic fusion processing on social media data, player behavior data, and game purchase records based on the fusion vector graph to obtain a fusion feature matrix; The fused feature matrix, social media data, player behavior data, and game purchase records are input into the deep neural network model to obtain the survey sample dataset.

2. The method for intelligent matching of dynamic market research samples based on multi-source data fusion according to claim 1 is characterized in that: Methods for analyzing player behavior data for activity and obtaining behavior pattern vectors include: Based on the behavior records of each player in the player behavior data, the timestamps in the behavior records are sorted to obtain a behavior time series, and feature extraction processing is performed on each time point in the behavior time series to obtain the behavior feature data of the player at different time points; Calculate the activity level of each player's behavioral feature data to obtain an activity score for each player; The activity scores of all players are normalized to obtain a normalized activity vector, and the normalized activity vectors of all players are clustered to obtain a behavior pattern vector.

3. The method for intelligent matching of dynamic market research samples based on multi-source data fusion according to claim 2 is characterized in that: The method for calculating the activity level of each player's behavioral characteristic data to obtain the activity level score of each player includes: Among them, H′ represents the activity score, n represents the number of behavioral feature data, and p i represents the probability distribution of the i-th behavior feature data, log(p i ) represents the logarithmic probability of the i-th behavioral characteristic data, α represents the timeliness coefficient, Δt i Represents the time difference of the i-th behavior feature data, 1+α·Δt i Used to adjust the weight of behavioral feature data based on time difference.

4. The method for intelligent matching of dynamic market research samples based on multi-source data fusion according to claim 2 is characterized in that: The method of performing multi-dimensional feature fusion processing on the emotion feature vector, the behavior pattern vector, and the consumption behavior vector to obtain a fusion vector graph includes: The emotional feature vector, behavioral pattern vector, and consumer behavior vector are mapped to the same dimensional space and normalized to obtain three types of normalized feature vectors with consistent structures. Each dimensional element in each normalized feature vector corresponds to a preset behavioral label index. Performing vector alignment fusion processing on the three types of normalized eigenvectors to construct an initial fusion matrix, wherein the vector alignment fusion processing refers to performing weighted splicing processing on the three types of eigenvectors in each dimension according to the index order; The initial fusion matrix is ​​reconstructed by feature density to obtain a fusion density matrix. The feature dimension corresponding to each behavior label index in the fusion density matrix is ​​used as a node, and the joint change intensity of any two feature dimensions in the initial fusion matrix is ​​used as an edge to construct a fusion vector graph.

5. The method for intelligent matching of dynamic market research samples based on multi-source data fusion according to claim 4 is characterized in that: Methods for performing feature density reconstruction on the initial fusion matrix to obtain a fusion density matrix include: Based on the feature dimension element value corresponding to each behavior label index in the initial fusion matrix, calculate its joint occurrence frequency in the three types of normalized feature vectors; All joint occurrence frequencies are organized into a vector structure according to the index order to obtain the original density vector, which is used to represent the joint response frequency of each feature dimension; Perform time series window sliding processing on the original density vector, evolve the sequence window according to the preset behavior label, perform local mean and range analysis on the density values ​​within each window range, and generate a set of density fluctuation factors; The original density vector is nonlinearly smoothed according to the density fluctuation factor set to generate a fused density matrix.

6. The method for intelligent matching of dynamic market research samples based on multi-source data fusion according to claim 5 is characterized in that: Methods for performing local mean and range analysis on density values ​​within each window range and generating a set of density fluctuation factors include: The density value sequence of each feature dimension in the fusion density matrix is ​​recorded as Q = {q1,q2,...,q L }, based on the window length u and step size v, Q is divided into sliding windows to obtain the density subsequence set R = {Q j }, where each Q j ={q j ,q j+1 ,...,q j+u-1 }, satisfying j∈{1,1+v,...,L-u+1}; For each Q j , calculate its second-order disturbance, based on the second-order disturbance, and introduce the exponential adjustment factor, construct the density fluctuation factor, and form all density fluctuation factors into a density fluctuation factor set.

7. The method for intelligent matching of dynamic market research samples based on multi-source data fusion according to claim 6 is characterized in that: For each Q j , the methods for calculating its second-order perturbation include: Among them, j represents the second-order perturbation corresponding to the jth window, u represents the length of the sliding window, u-1 represents the number of time points in each window that can be used to calculate the perturbation, and q j+k represents the density value at the j+kth time point, q j+k-1 Indicates the density value at the previous time point, (q j+k -q j+k-1 ) 2 represents the square of the difference between two moments, Indicates averaging of the sum of squared disturbances.

8. The method for intelligent matching of dynamic market research samples based on multi-source data fusion according to claim 6 is characterized in that: Based on the second-order perturbation and introducing the exponential adjustment factor, the method of constructing the density fluctuation factor includes: Among them, θ j represents the density fluctuation factor corresponding to the j-th window, represents the cumulative result of all squares of density values ​​in the jth window, ∈ represents the smoothing constant, exp(-λ·σ j ) represents the standard deviation of the disturbance σ j The exponential adjustment factor constructed, where λ represents the adjustment coefficient, σ j represents the standard deviation of the density value in the j-th window, 9. The method for intelligent matching of dynamic market research samples based on multi-source data fusion according to claim 6 is characterized in that: Methods for performing weighted dynamic fusion processing on social media data, player behavior data, and game purchase records based on the fusion vector graph to obtain a fusion feature matrix include: Obtain the node set and edge set in the fusion vector graph, where the node set represents the feature dimension corresponding to each line label index, and the edge set represents the joint change strength between any two feature dimensions; Based on the behavioral label index order of each type of original feature vector in social media data, player behavior data, and game purchase records, the three types of original feature vectors are mapped to the node set of the fusion vector graph to generate three types of index-bound feature vectors; Normalize the weights of the edge sets in the fusion vector graph to obtain the normalized edge weight vector. Perform dot product processing on the normalized edge weight vector and the corresponding elements in the three types of index binding feature vectors to generate the fusion response vector. Perform interval segmentation processing on the fusion response vector, map the corresponding index to the fusion level label according to the interval to which the response value belongs, and combine all indexes in the level label order to generate a level label set; A fusion feature matrix is ​​constructed based on the level label set and the fusion response value corresponding to each index in the fusion response vector.

10. The method for intelligent matching of dynamic market research samples based on multi-source data fusion according to claim 9, characterized in that: The method of performing dot product processing on the normalized edge weight vector and the corresponding elements in the three-category index binding feature vector to generate a fused response vector includes: According to the frequency of occurrence of the feature dimension connected by each edge in the normalized initial vector in the three types of index-bound feature vectors, a set of dynamic adjustment factors is constructed, where each dynamic adjustment factor is used to reflect the local sensitivity of the feature dimension to the change of the response value; Perform element-wise product processing on the normalized edge weight initial vector and the adjustment factor at the corresponding position in the dynamic adjustment factor set to generate a fused edge weight vector; The fusion response vector is constructed by performing weighted summation processing on the fusion edge weight vector and the values ​​of the corresponding feature dimensions in the three-category index binding feature vector; According to the order of weighted values ​​corresponding to all behavior label indexes in the fusion response vector, the column structure of the fusion feature matrix is ​​organized and generated to obtain the fusion feature matrix.