A brand information analysis method based on big data
By generating a multi-source time stamp set and calculating the click path offset, combined with the text overlap parameter, the problem of data correlation breakage in multi-channel brand information analysis was solved, thereby improving the accuracy and continuity of brand information analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN AGRI UNIV
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies lack cross-analysis of text data from multiple channels within the same time frame when processing brand information analysis. This results in the inability to record differences in keyword change nodes and path selection, affecting the ability to identify the relationship between behavior-driven and content-response relationships.
By generating a multi-source time stamp set, identifying keyword change nodes and calculating click path offsets, and combining text overlap parameters, the system integrates user path changes with brand text consistency to form a joint feature vector, thereby achieving time alignment and correlation analysis of multi-channel data.
It enhances the ability to express the correlation between behavioral changes and content changes, avoids the problem of broken correlation caused by independent calculation of a single data dimension, and enhances the accuracy and continuity of brand information analysis.
Smart Images

Figure CN122491271A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of brand analysis technology, and in particular to a brand information analysis method based on big data. Background Technology
[0002] The field of brand analytics technology encompasses a collection of technologies based on computer information processing and business data analysis. It involves core aspects such as data acquisition, data cleaning, data structuring, modeling, and the expression of analysis results. This field revolves around user behavior records, transaction data, textual information, and online interaction data. It constructs a feature representation system through data correlation and combines statistical analysis methods, time series analysis methods, and classification and clustering methods to identify, organize, and analyze brand-related information, thereby forming a systematic information processing framework oriented towards the market environment.
[0003] Among them, the brand information analysis method based on big data refers to a specific method for integrating, processing, and analyzing multi-source brand-related data. The brand-related data includes product review data from e-commerce platforms, social media text data, search engine keyword data, and user click behavior records. This method covers the following stages: data collection stage, which obtains structured and unstructured data through interfaces; data preprocessing stage, which performs word segmentation and stop word removal on the text and constructs word frequency statistics; feature construction stage, which forms a feature set based on word frequency statistics and word vector representation; and analysis stage, which uses clustering methods to classify brand sentiment and uses time series methods to model changes in brand attention. At the same time, it uses a sentiment dictionary to annotate the sentiment tendency of the text, thereby constructing the brand information analysis process.
[0004] Current processing methods focus on word frequency statistics and text expression construction. Keyword input sequences and click path records are mostly processed as independent structures, lacking time tag alignment and path branching relationship characterization. Text data processing is concentrated on term statistics and sentiment annotation, without performing cross-analysis on text time intervals from different sources. This results in the lack of content relationships across multiple channels within the same time range. For example, when users adjust keywords and cause page jumps during the search process, the differences in keyword change nodes and path selections cannot be recorded, and behavioral change information is difficult to reflect in feature expression. At the same time, when texts from multiple platforms publish related content in the same time period, content consistency is not quantified, resulting in a discrete state in the brand information dissemination process, affecting the ability to identify the relationship between behavioral drivers and content responses in subsequent analysis. Summary of the Invention
[0005] The main objective of this invention is to provide a brand information analysis method based on big data, which can effectively solve the problems of unreasonable spatial layout of mobile blood purification treatment and poor overall medical service level and emergency response capability.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A brand information analysis method based on big data includes the following steps: S1. Data Acquisition and Preprocessing: Obtain records of keyword input and click paths for product reviews posted on social media, extract timestamp text rating page identifier fields and sort them by time to generate a multi-source timestamp set; S2. Keyword Change Identification: Extract keyword sequences based on multi-source time stamp sets, perform adjacent character comparison to count the number and position of replacements, compare the number of replacements with a threshold and mark the change nodes to generate a keyword change identifier sequence; S3. Click path branch offset calculation: Extract the source page identifier, target page identifier, and jump time of the click path based on the keyword change identifier sequence. Form a branch set for the same source page and mark the path number. Calculate the difference between the change node time and the jump time to filter records. Extract the path number set before and after the change, perform the difference operation, and calculate the ratio with the total number of paths to generate a path offset coefficient sequence. S4. Calculation of text overlap at time intersections: Extract text time intervals based on multi-source time stamp sets and determine the overlap of intervals. Number and mark text intervals with time overlaps. Perform word segmentation on the text within the overlapping intervals and generate a term set. Count the overlapping terms between texts through intersection operations. Combine the total number of terms to complete the normalization process and clearly present the semantic fit of brand texts from multiple channels within the same time range. S5. Comprehensive Brand Information Correlation Analysis: Retrieve path offset analysis results and semantic fit of multi-channel texts, and complete data matching by time tag and keyword number; perform coupling operation and time sequence sorting on the matched data, integrate the inherent relationship between user path changes and brand text consistency, and form information analysis conclusions that can be directly used for brand assessment.
[0007] Preferably, the multi-source time stamp set includes time-series behavior records, structured field groups, and normalized index items; the keyword change identifier sequence includes mutation node markers, trend classification codes, and sequence state values; the path offset coefficient sequence includes branch difference rate, jump correlation degree, and time-series deviation value; the text overlap parameter includes cross-interval matching degree, term overlap rate, and semantic similarity; and the joint feature numerical vector includes time-series correlation value, weighted product term, and sorting feature value.
[0008] Preferably, the data acquisition and preprocessing includes the following steps: S101. Obtain product reviews, social media posts, keyword input commands, and click path records. Extract text ratings and text content from product reviews, extract keyword items from social media posts, extract page identifier fields from click path records, and timestamps corresponding to each operation to establish a basic dimension data information set. S102. Call the basic dimension data information set, associate and match the text score with the page identifier field, arrange the operation records and evaluation content in chronological order according to the timestamp, calculate the time interval difference between each operation step, compare with the preset logical jump benchmark value, and generate the chronological association logic quantity. S103. Based on the time-series correlation logic quantity, perform multi-source normalization operation on the keyword items, text scores and page identifier fields in the basic dimension data information set, cluster and superimpose the text scores and keyword items under the same page identifier, obtain multi-source marked items with time attributes, and establish a multi-source time mark set.
[0009] Preferably, the keyword variation identification includes the following steps: S201. Extract keyword sequences based on multi-source time stamp sets, obtain the content of adjacent characters in the keyword sequences, compare the code positions of two adjacent characters, count the number of characters with inconsistent character codes, record the index positions of inconsistent characters, calculate the edit distance between adjacent strings, and obtain the character replacement difference. S202. Call the character replacement difference quantity, extract the replacement quantity and position information from the difference quantity, perform a numerical comparison logic operation based on the replacement quantity and the preset change quantity threshold, determine the position where the value exceeds the threshold range, perform logical setting and assignment on the corresponding sequence coordinates, and generate a keyword change identifier sequence.
[0010] Preferably, the click path branch offset calculation includes the following steps: S301. Extract the source page identifier, target page identifier, and jump time of the click path based on the keyword change identifier sequence. Logically merge the same source page identifiers to form a source path branch set. Assign a unique index value to each independent path in the set and establish the path branch index. S302. Call the keyword change identifier sequence and path branch index, extract the change node time and page jump time, perform subtraction to calculate the absolute difference between the two time nodes, compare the absolute difference with the preset path matching time threshold, retain the difference within the threshold range and obtain the effective jump spatiotemporal feature value. S303. Based on the effective jump spatiotemporal feature value, extract the path number set before and after the change, perform difference operation, count the number of path identifiers in the difference set, perform a division operation between the number and the total number of paths to obtain the ratio result, and map and arrange them according to the time axis sequence to generate the path offset coefficient sequence.
[0011] Preferably, the calculation of the time-intersecting text overlap includes the following steps: S401. Extract the text time interval based on the multi-source time stamp set, call the start time point and end time point of each text to perform logical overlap judgment, extract the text pairs that meet the time intersection conditions and record the intersection interval number, perform word segmentation processing on the text within the intersection interval, and obtain a term mapping set composed of multiple independent word units. S402. Based on the term mapping set, extract the term content under different text sequences, perform intersection operation, count the number of overlapping terms, call the total number of terms as the denominator, divide the number of overlapping terms by it, calculate the numerical overlap ratio, generate text overlap parameter, and complete normalization processing in combination with the total number of terms to clearly present the semantic fit of brand texts from multiple channels within the same time range.
[0012] Preferably, the construction of the joint feature vector includes the following steps: S501. Call the path offset coefficient sequence and text overlap parameter, extract the time label and keyword number corresponding to each data, execute the matching judgment logic of the same time axis coordinate and number, and perform the product operation on the offset coefficient and overlap value that pass the matching verification to obtain the multidimensional mapping product that reflects the spatiotemporal correlation strength. S502. Based on the multidimensional mapping product, extract the product values of each group and sort them in ascending order according to the time tag sequence. Fill the sorted values into the preset dimension vector matrix in order, perform vector space coordinate mapping, establish a joint feature value vector that conforms to the time sequence distribution characteristics, integrate the inherent relationship between user path changes and brand text consistency, and form information analysis conclusions that can be directly used for brand judgment.
[0013] Compared with the prior art, the present invention has the following beneficial effects: A unified time stamp is formed by recording the text time intervals around the click path jump of the keyword input sequence. The number of character replacements and the position index are used to mark the keyword change nodes and are aligned and filtered with the click path time difference. The source page identifier and the target page identifier constitute a branch set. The difference set is calculated as a ratio to the total number of paths to form the path offset. At the same time, cross-judgment is performed on the time intervals of text from different sources. The number of overlapping terms is calculated as a ratio to the total number of terms to form the text overlap parameter. Further, matching is performed based on the time tag and keyword number. The path offset and text overlap are multiplied and sorted to form a joint feature value vector. This establishes a correspondence between the keyword change behavior click path structure change and the text content from multiple sources in the same time dimension. The data forms a continuous mapping chain, which improves the ability to express the correlation between behavior change and content change and avoids the problem of broken correlation caused by independent calculation of a single data dimension. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the overall process of the present invention. Detailed Implementation
[0015] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0016] Example 1, as Figure 1 As shown, a brand information analysis method based on big data includes the following steps: S1. Data Acquisition and Preprocessing: Obtain records of keyword input and click paths for product reviews posted on social media, extract timestamp text rating page identifier fields and sort them by time to generate a multi-source timestamp set.
[0017] S2. Keyword Change Identification: Extract keyword sequences based on multi-source time stamp sets, perform adjacent character comparison to count the number and position of replacements, compare the number of replacements with a threshold and mark the change nodes to generate a keyword change identifier sequence; S3. Click path branch offset calculation: Extract the source page identifier, target page identifier, and jump time of the click path based on the keyword change identifier sequence. Form a branch set for the same source page and mark the path number. Calculate the difference between the change node time and the jump time to filter records. Extract the path number set before and after the change, perform the difference operation, and calculate the ratio with the total number of paths to generate a path offset coefficient sequence. S4. Calculation of text overlap at time intersections: Extract text time intervals based on multi-source time stamp sets and determine the overlap of intervals. Number and mark text intervals with time overlaps. Perform word segmentation on the text within the overlapping intervals and generate a term set. Count the overlapping terms between texts through intersection operations. Combine the total number of terms to complete the normalization process and clearly present the semantic fit of brand texts from multiple channels within the same time range. S5. Comprehensive Brand Information Correlation Analysis: Retrieve the path offset analysis results and the semantic fit of multi-channel texts, and complete data matching by time tag and keyword number; perform coupling operation and time sequence sorting on the matched data, integrate the inherent relationship between user path changes and brand text consistency, and form information analysis conclusions that can be directly used for brand assessment. The multi-source time stamp set includes time-series behavior records, structured field groups, and normalized index terms; the keyword change identifier sequence includes mutation node markers, trend classification codes, and sequence state values; the path offset coefficient sequence includes branch difference rate, jump correlation degree, and time-series deviation value; the text overlap parameter includes cross-interval matching degree, term overlap rate, and semantic similarity; and the joint feature numerical vector includes time-series correlation value, weighted product term, and ranking feature value.
[0018] Data acquisition and preprocessing includes the following steps: S101. Obtain product reviews, social media posts, keyword input commands, and click path records. Extract text ratings and text content from product reviews, extract keywords from social media posts, and extract page identifier fields and timestamps corresponding to each operation from click path records. Establish a basic dimensional data set, including product review data, social media posts, user behavior click path records, and keyword input commands. 1. Extract the text rating field and review text content from product review data; 2. Extract a set of keywords from social media posts; 3. Extract the page identifier field and the timestamp corresponding to each operation from the click path record; The above data is structured to construct a set of basic dimensional data information including "text rating, keyword items, page identifier fields, and timestamps". S102. Call the basic dimension data information set, associate and match the text score with the page identifier field, arrange the operation records and evaluation content in chronological order according to the timestamp, calculate the time interval difference between each operation step, compare with the preset logical jump benchmark value, generate the chronological association logic quantity, call the basic dimension data information set, and associate and map the text score with the page identifier field. User actions and evaluation content are uniformly sorted chronologically based on timestamps; Calculate the time interval difference Δt between adjacent operations; The time interval difference Δt is compared with the preset logical jump baseline value T0 to generate a temporal correlation logic quantity L that represents the continuity of user behavior. t ; Among them, the time-related logic quantity L t It should include at least: continuous access indicator, abnormal redirection indicator, and dwell time weight value; S103. Based on the time-series correlation logic quantity, perform multi-source normalization operation on the keyword items, text scores and page identifier fields in the basic dimension data information set, cluster and superimpose the text scores and keyword items under the same page identifier, obtain multi-source labeled items with time attributes, and establish a multi-source time label set. Based on the time-related logical quantity L t Multi-source normalization is performed on the keyword items, text scores, and page identifier fields in the basic dimension data information set; Grouping according to the page identifier field, and clustering and fusing text scores and keyword items under the same page identifier to form related feature clusters; By combining timestamp information, time attribute labels are added to the associated feature clusters to generate multi-source labeled entries with temporal features; All multi-source marker entries are summarized to construct a multi-source time stamp set.
[0019] Keyword variation identification includes the following steps: S201. Extract keyword sequences based on multi-source time stamp sets, obtain the content of adjacent characters in the keyword sequences, compare the code positions of two adjacent characters, count the number of characters with inconsistent character codes, record the index positions of inconsistent characters, calculate the edit distance between adjacent strings, and obtain the character replacement difference. Based on a multi-source time stamp set, extract keyword sequences with time attributes; Arrange the keyword sequence in chronological order and obtain adjacent keyword string pairs; For each pair of adjacent keyword strings, perform the following processing: 1. Perform character-level splitting on the string to obtain the character sequence; 2. Obtain the encoding value corresponding to each character; 3. Compare the encoding values of adjacent characters one by one, count the number of characters with inconsistent encodings, and record the set of index positions of inconsistent characters; 4. Calculate the edit distance value D between adjacent keyword strings based on the character sequence. e ; Compare the coding difference statistics with the edit distance value D eThe components are merged to generate a character substitution difference Δc that represents the degree of keyword change. Δc = Coding difference + Edit distance fusion amount The character replacement difference Δc includes at least the number of characters replaced, the set of replacement position indices, and the edit distance weight value.
[0020] S202. Call the character replacement difference quantity, extract the replacement quantity and position information from the difference quantity, perform a numerical comparison logic operation based on the replacement quantity and the preset change quantity threshold, determine the position where the value exceeds the threshold range, perform logical setting and assignment on the corresponding sequence coordinates, and generate a keyword change identifier sequence. Call the character replacement difference Δc to extract the number of replaced characters and their corresponding index positions; Compare the number of replaced characters with the preset change threshold N0, and execute the threshold determination logic: When the number of replaced characters is greater than or equal to N0, it is determined to be a position with significant changes; When the number of replaced characters is less than N0, it is determined to be a position with insignificant change; For positions identified as having significant changes, a logical bit-setting operation is performed at the index coordinates of the corresponding keyword sequence to generate a binary change identifier; Perform logical bit-setting on all keyword sequences to construct a keyword change identifier sequence; Among them, the keyword change identifier sequence is used to characterize the dynamic change features of keywords in the time dimension; Based on keyword change identifier sequences and multi-source time stamp sets, a comprehensive analysis of keyword changes over time is conducted. Key positions of continuous or abrupt changes in the keyword change identifier sequences are identified. By combining the corresponding timestamp information, abnormal mutation nodes and change trend paths in the keyword evolution process are determined. High-frequency evolution patterns are extracted based on change frequency and distribution characteristics, thereby generating analysis results to characterize the dynamic evolution features of keywords.
[0021] The path branch offset calculation involves the following steps: S301. Extract the source page identifier, target page identifier, and jump time of the click path based on the keyword change identifier sequence. Logically merge the same source page identifiers to form a source path branch set. Assign a unique index value to each independent path in the set and establish the path branch index. Based on the keyword change identifier sequence, extract the source page identifier, target page identifier, and corresponding jump time from the user click path; Based on the source page identifier, click paths with the same source page identifier are logically aggregated to form a source path branch set; Each independent jump path in the source path branch set is numbered and assigned a unique index identifier to construct a path branch index.
[0022] S302. Call the keyword change identifier sequence and path branch index, extract the change node time and page jump time, perform subtraction to calculate the absolute difference between the two time nodes, compare the absolute difference with the preset path matching time threshold, retain the difference within the threshold range and obtain the effective jump spatiotemporal feature value. Call the keyword change identifier sequence and path branch index to extract the time corresponding to the keyword change node and the page jump time; Calculate the difference between the keyword change time and the page jump time, and obtain the absolute time difference Δt. p ; The absolute time difference Δt p Matching time threshold T with preset path p Compare and select those that satisfy Δt p ≤T p Path records; Summarize the screening results and obtain effective spatiotemporal feature values that characterize the relationship between keyword changes and path jumps; S303. Based on the effective jump spatiotemporal feature values, extract the path number sets before and after the change, perform difference set operation, count the number of path identifiers in the difference set, perform a division operation between the number and the total number of paths to obtain the ratio result, and map and arrange them according to the time axis sequence to generate a path offset coefficient sequence. Based on the effective spatiotemporal feature values of the jump, the set of path numbers before and after the keyword change is extracted respectively; Perform a difference operation on the set of path numbers to obtain the set of path differences before and after the change; The number of path identifiers in the statistical difference set is counted and normalized with the total number of corresponding paths to calculate the path offset coefficient. The path offset coefficients are mapped and arranged according to the time series to generate a path offset coefficient sequence; Among them, the path offset coefficient is used to characterize the degree of influence of keyword changes on the distribution of user access paths; Based on the path offset coefficient sequence, abnormal path offset intervals are identified, and user behavior offset analysis results are generated by combining keyword change characteristics. These results can be used to identify changes in brand attention or make page optimization decisions.
[0023] The calculation of text overlap over time includes the following steps: S401. Extract the text time interval based on the multi-source time stamp set, call the start time point and end time point of each text to perform logical overlap judgment, extract the text pairs that meet the time intersection conditions and record the intersection interval number, perform word segmentation processing on the text within the intersection interval, and obtain a term mapping set composed of multiple independent word units. Based on the multi-source time stamp set, extract the start and end time points corresponding to each text data to construct the text time interval; Perform overlap determination operation on each text time interval, filter text pairs that meet the time intersection condition, and assign a unique number to the corresponding intersection time interval; For each cross-time interval, perform word segmentation on the text data to obtain a set of terms consisting of multiple independent word units; Perform unified mapping encoding on the term set to construct a term mapping set. S402. Based on the term mapping set, extract the term content under different text sequences, perform intersection operation, count the number of overlapping terms, call the total number of terms as the denominator, divide the number of overlapping terms with it, calculate the numerical overlap ratio, generate text overlap parameter, and complete the normalization process in combination with the total number of terms to clearly present the semantic fit of brand texts from multiple channels within the same time range. Based on the term mapping set, extract the term set corresponding to different text sequences; Perform an intersection operation on the term sets to obtain the set of overlapping terms, and count the number N of overlapping terms. o ; Get the total number N of terms involved in the calculation. t (Preferably, the number of terms in the union set or the weighted total); The number of overlapping terms N o Total number of terms N t Perform normalization calculations to obtain the text overlap parameter R. t ; Among them, the text overlap parameter R t Used to characterize the degree of semantic consistency between different text contents within a time crossover interval, the total number of terms is the weighted number of terms, and the weight is determined based on the keyword change identifier sequence or the frequency of term occurrence.
[0024] The construction of joint feature vectors involves the following steps: S501. Call the path offset coefficient sequence and text overlap parameter, extract the time label and keyword number corresponding to each data, execute the matching judgment logic of the same time axis coordinate and number, and perform the product operation on the offset coefficient and overlap value that pass the matching verification to obtain the multidimensional mapping product that reflects the spatiotemporal correlation strength. Call the path offset coefficient sequence and text overlap parameter to extract the time tag and keyword number corresponding to each data point; Based on time tags and keyword numbers, a matching judgment is performed to filter out data pairs that meet the conditions of the same time axis coordinate and the same keyword number. For successfully matched data pairs, a coupling operation is performed on the path offset coefficient and the text overlap parameter to obtain a multidimensional mapping product that represents the degree of spatiotemporal correlation. The coupling operation is either a product operation or a weighted product operation with weight coefficients, used to enhance the expression of the correlation strength between different feature dimensions. S502. Based on the multidimensional mapping product, extract the product values of each group and sort them in ascending order according to the time tag sequence. Fill the sorted values into the preset dimension vector matrix in order, perform vector space coordinate mapping, establish a joint feature value vector that conforms to the time sequence distribution characteristics, integrate the inherent relationship between user path changes and brand text consistency, and form information analysis conclusions that can be directly used for brand judgment. Based on the multidimensional mapping product, extract the product values of each group and sort them in ascending order according to the time label; The sorted values are mapped sequentially into a vector matrix of a preset dimension to construct a time-ordered sequence of feature values. Perform vector space mapping on the feature numerical sequence to generate a joint feature numerical vector with time distribution characteristics; Among them, the joint feature numerical vector is used to characterize the comprehensive correlation features between keyword variation, path offset and text consistency.
[0025] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A brand information analysis method based on big data, comprising the following steps: S1. Data Acquisition and Preprocessing: Obtain records of keyword input and click paths for product reviews posted on social media, extract timestamp text rating page identifier fields and sort them by time to generate a multi-source timestamp set; S2. Keyword Change Identification: Extract keyword sequences based on multi-source time stamp sets, perform adjacent character comparison to count the number and position of replacements, compare the number of replacements with a threshold and mark the change nodes to generate a keyword change identifier sequence; S3. Click path branch offset calculation: Extract the source page identifier, target page identifier, and jump time of the click path based on the keyword change identifier sequence. Form a branch set for the same source page and mark the path number. Calculate the difference between the change node time and the jump time to filter records. Extract the path number set before and after the change, perform the difference operation, and calculate the ratio with the total number of paths to generate a path offset coefficient sequence. S4. Calculation of text overlap at time intersections: Extract text time intervals based on multi-source time stamp sets and determine the overlap of intervals. Number and mark text intervals with time overlaps. Perform word segmentation on the text within the overlapping intervals and generate a term set. Count the overlapping terms between texts through intersection operations. Combine the total number of terms to complete the normalization process and clearly present the semantic fit of brand texts from multiple channels within the same time range. S5. Comprehensive Brand Information Correlation Analysis: Retrieve path offset analysis results and semantic fit of multi-channel texts, complete data matching by time tag and keyword number, perform coupling operation and time sequence sorting on the matched data, integrate the inherent relationship between user path changes and brand text consistency, and form information analysis conclusions that can be directly used for brand assessment.
2. The brand information analysis method based on big data according to claim 1, characterized in that: The multi-source time stamp set includes time-series behavior records, structured field groups, and normalized index items; the keyword change identifier sequence includes mutation node markers, trend classification codes, and sequence state values; the path offset coefficient sequence includes branch difference rate, jump correlation degree, and time-series deviation value; the text overlap parameter includes cross-interval matching degree, term overlap rate, and semantic similarity; and the joint feature numerical vector includes time-series correlation value, weighted product term, and sorting feature value.
3. The brand information analysis method based on big data according to claim 1, characterized in that, The data acquisition and preprocessing includes the following steps: S101. Obtain product reviews, social media posts, keyword input commands, and click path records. Extract text ratings and text content from product reviews, extract keyword items from social media posts, extract page identifier fields from click path records, and timestamps corresponding to each operation to establish a basic dimension data information set. S102. Call the basic dimension data information set, associate and match the text score with the page identifier field, arrange the operation records and evaluation content in chronological order according to the timestamp, calculate the time interval difference between each operation step, compare with the preset logical jump benchmark value, and generate the chronological association logic quantity. S103. Based on the time-series correlation logic quantity, perform multi-source normalization operation on the keyword items, text scores and page identifier fields in the basic dimension data information set, cluster and superimpose the text scores and keyword items under the same page identifier, obtain multi-source marked items with time attributes, and establish a multi-source time mark set.
4. The brand information analysis method based on big data according to claim 1, characterized in that, The keyword variation identification includes the following steps: S201. Extract keyword sequences based on multi-source time stamp sets, obtain the content of adjacent characters in the keyword sequences, compare the code positions of two adjacent characters, count the number of characters with inconsistent character codes, record the index positions of inconsistent characters, calculate the edit distance between adjacent strings, and obtain the character replacement difference. S202. Call the character replacement difference quantity, extract the replacement quantity and position information from the difference quantity, perform a numerical comparison logic operation based on the replacement quantity and the preset change quantity threshold, determine the position where the value exceeds the threshold range, perform logical setting and assignment on the corresponding sequence coordinates, and generate a keyword change identifier sequence.
5. The brand information analysis method based on big data according to claim 1, characterized in that, The calculation of the click path branch offset includes the following steps: S301. Extract the source page identifier, target page identifier, and jump time of the click path based on the keyword change identifier sequence. Logically merge the same source page identifiers to form a source path branch set. Assign a unique index value to each independent path in the set and establish the path branch index. S302. Call the keyword change identifier sequence and path branch index, extract the change node time and page jump time, perform subtraction to calculate the absolute difference between the two time nodes, compare the absolute difference with the preset path matching time threshold, retain the difference within the threshold range and obtain the effective jump spatiotemporal feature value. S303. Based on the effective jump spatiotemporal feature value, extract the path number set before and after the change, perform difference operation, count the number of path identifiers in the difference set, perform a division operation between the number and the total number of paths to obtain the ratio result, and map and arrange them according to the time axis sequence to generate the path offset coefficient sequence.
6. The brand information analysis method based on big data according to claim 1, characterized in that, The calculation of text overlap over time includes the following steps: S401. Extract the text time interval based on the multi-source time stamp set, call the start time point and end time point of each text to perform logical overlap judgment, extract the text pairs that meet the time intersection conditions and record the intersection interval number, perform word segmentation processing on the text within the intersection interval, and obtain a term mapping set composed of multiple independent word units. S402. Based on the term mapping set, extract the term content under different text sequences, perform intersection operation, count the number of overlapping terms, call the total number of terms as the denominator, divide the number of overlapping terms by it, calculate the numerical overlap ratio, generate text overlap parameter, and complete normalization processing in combination with the total number of terms to clearly present the semantic fit of brand texts from multiple channels within the same time range.
7. The brand information analysis method based on big data according to claim 1, characterized in that, The construction of the joint feature vector includes the following steps: S501. Call the path offset coefficient sequence and text overlap parameter, extract the time label and keyword number corresponding to each data, execute the matching judgment logic of the same time axis coordinate and number, and perform the product operation on the offset coefficient and overlap value that pass the matching verification to obtain the multidimensional mapping product that reflects the spatiotemporal correlation strength. S502. Based on the multidimensional mapping product, extract the product values of each group and sort them in ascending order according to the time tag sequence. Fill the sorted values into the preset dimension vector matrix in order, perform vector space coordinate mapping, establish a joint feature value vector that conforms to the time sequence distribution characteristics, integrate the inherent relationship between user path changes and brand text consistency, and form information analysis conclusions that can be directly used for brand judgment.