Method and System for Generating Technical Barriers to Trade Questionnaire Based on Big Data

By analyzing the information text and web traffic data of the trade service platform, building a problem sequence model, and generating a technical trade measure questionnaire, the problems of low efficiency and poor flexibility in the existing methods are solved, and the scientificity and timeliness of questionnaire generation are improved.

CN118886951BActive Publication Date: 2025-07-04CHINA NAT INST OF STANDARDIZATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410954920.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2025-07-04
Estimated Expiration
2044-07-17

AI Technical Summary

Technical Problem

The existing method of questionnaire generation of technical trade measures relies on human subjective judgment, has low data collection and processing efficiency, lacks attention to user-focused content, and lacks flexibility and scalability, making it difficult to adapt to market changes.

Method used

By obtaining the information text data and web page traffic data of the trade service platform, after preprocessing, the TF-IDF algorithm is used to extract the semantic importance of keywords and user interest, construct a bias analysis function, construct a problem sequence model, calculate the contribution of the problem based on the simulation results, and generate a questionnaire.

Benefits of technology

It improves the efficiency and scientificity of the questionnaire generation, can flexibly adjust the questionnaire content, capture trade hot issues, reflect users' potential preferences, and improves the timeliness and market relevance of the questionnaire.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118886951B_ABST
    Figure CN118886951B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for generating a technical trade measure questionnaire based on big data, including obtaining and preprocessing information text data and web traffic data of a trade service platform, performing keyword analysis on the information text data to obtain keyword semantic importance, performing interest analysis on the web traffic data to obtain user interest, constructing a bias analysis function based on the keyword semantic importance and user interest, constructing a question sequence model, embedding the bias analysis function into the question sequence model, simulating the question sequence model, calculating the contribution degree of questions based on the simulation results, determining a questionnaire generation scheme according to the contribution degree of questions, and generating a questionnaire based on the questionnaire generation scheme. This method can not only improve the generation efficiency and quality of the questionnaire, but also has good interpretability and can be directly applied to the technical trade measure questionnaire evaluation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of questionnaire generation, and in particular to a method and system for generating technical trade measure questionnaires based on big data. Background Art

[0002] With the continuous deepening of globalized trade, technical trade measures have an increasingly significant impact on international trade, and are also closely related to national security and economic trends, which in turn leads to greater attention from consumers and the general public to relevant information such as technical trade measures. Therefore, the investigation and analysis of technical trade measures are of great significance for enterprises to formulate export strategies and governments to formulate trade policies.

[0003] Traditional questionnaire generation methods mainly rely on subjective human judgment, but this method has limitations, such as low data collection and processing efficiency, and a lack of objective judgment on trade measures. In recent years, with the rapid development of big data technology, more and more fields have begun to use big data methods to generate questionnaires to improve data collection and processing efficiency.

[0004] At present, although there are already some technical trade measure investigation and analysis methods based on big data, these methods often only focus on the question generation method, lack attention to the content emphasized by users, and lack research on the generation order of questions in the questionnaire. In addition, existing questionnaire generation methods and systems often lack flexibility and scalability, and are difficult to adapt to the changing trade environment in the market. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for generating technical trade measure questionnaires based on big data.

[0006] To achieve the above object, the present invention is implemented according to the following technical solutions:

[0007] The first aspect of the present invention provides a method for generating technical trade measure questionnaires based on big data, including the following steps:

[0008] S1 Obtain the information text data and web traffic data of the trade service platform, and perform preprocessing;

[0009] S2 Perform keyword analysis on the information text data to obtain the semantic importance of keywords, and perform interest analysis on the web traffic data to obtain user interest;

[0010] S3 Construct a bias analysis function based on the semantic importance of keywords and the user interest;

[0011] S4 Construct a question sequence model, and embed the bias analysis function into the question sequence model;

[0012] S5 simulates the problem sequence model, calculates the contribution degree of the problem based on the simulation result, and generates a questionnaire according to the contribution degree.

[0013] Further, the method for obtaining the keyword semantic importance by performing keyword analysis on the information text data in step S2 includes:

[0014] S11 extracts all words in the information text data through the TF-IDF algorithm;

[0015] S12 constructs a keyword graph structure for all words. Each node in the graph represents a word, and the edges between nodes represent the semantic relationship between words. The weight of each node is obtained through iterative calculation. The iterative calculation formula is:

[0016]

[0017] Among them, ω(V i ) is the weight value of node V i , d is the damping coefficient, In(V i ) represents the set of all nodes pointing to node V i , w ji represents the weight value of the edge from node V j to node V i , Out(V j ) represents all nodes pointed to by node V j , w jk represents the weight value of the edge from node V j to node V k , ω(V j ) is the weight value of node j. Sort in descending order according to the node weight value, and use the word corresponding to the node with the highest weight as the final keyword;

[0018] S13 calculates the semantic importance of the keyword. The calculation formula is:

[0019]

[0020] Among them, l is the length of keyword m, is the average character length of the keyword, |D| represents the total number of words in the text, N is the number of keywords, C k represents the sum of the occurrence frequencies of all keywords except keyword m in the set {1, 2,... N}, C m is the occurrence frequency of keyword m, ω(V m ) represents the node weight value corresponding to keyword m, ‖ω(V m )‖ represents the Euclidean norm of the node weight value, |{j:t i∈d j}| represents the number of statements in the text that include the keyword m, I m is the inverse document frequency of the statements that include the keyword m, represents the sum of the inverse document frequencies of all keywords except the keyword m in the set {1, 2,... N}.

[0021] Further, the method for obtaining the user interest degree by analyzing the interest degree of the web traffic data includes:

[0022] S21 Obtain the http request log frame and the click event frame to extract key data to obtain the user click data of the information, and obtain the page view time frame and the scroll depth frame to obtain the user information browsing data;

[0023] S22 Introduce a time decay factor, and calculate the user interest preference degree based on the user click data and the user information browsing data. The calculation formula is:

[0024]

[0025] λ n =(t s -5)×0.1

[0026] where n is the number of paragraphs of the information text data, is the average click-through rate of the user, is the average browsing time of the user, is expressed as the probability after normalizing the user action vector, μ n is the fatigue factor, which represents the reading fatigue value of the user for the reading content as time increases, λ n is the time decay factor, where (t s -5)×0.1 is expressed as a penalty term, indicating that when the browsing time of the user for the information is less than 5s, the user's interest preference degree is reduced.

[0027] Further, the method for constructing a bias analysis function based on the semantic importance of the keyword and the user interest degree includes:

[0028] Construct a bias analysis function based on the user interest degree and the semantic importance corresponding to the selected keywords. The expression is:

[0029]

[0030] where k is the number of keywords whose semantic importance is greater than the average semantic importance, k j and k t are the fuzzy linguistic values of the keyword and the statement corresponding to the keyword respectively, m is the central scale value of the wavelet function, f is the wavelet function, m ij and σij are the translation parameter and dilation parameter of the wavelet function respectively, and score j is the semantic importance of keywords whose semantic importance is greater than the average semantic importance, score2 is the user interest degree, and T is the number of all keywords. is the average node weight value of all keywords. means that only keywords whose semantic importance is greater than the average semantic importance are analyzed.

[0031] Furthermore, the method for constructing the problem sequence model includes:

[0032] S31 constructs a first problem model and designs a first network structure layer, including: an input layer, an embedding vector layer, an encoding layer, a fully connected layer, a Softmax layer, and an output layer. The processing of data is to input the input sequence as input data into the embedding vector layer, perform encoding processing through the encoding layer, and then complete problem generation through the fully connected layer and the Softmax layer. Among them, the vector output by the encoder with a self-attention mechanism obtains the corresponding problem output through the fully connected layer. The expressions of the encoding and decoding methods are:

[0033] I input =[CLS],Q1,Q2,...,Q m ,[SEP],C1,C2,...,C m ,[SEP]

[0034]

[0035] P(Q i )=Softmax(f(Q i-1 ))

[0036] Among them, I input is the input data, [CLS],Q1,Q2,...,Q m represents the input sequence, m is the number of input sequences, [SEP],C1,C2,...,C n ,[SEP] represents the target problem output sequence, f(Q i ) is the vector of the i-th element Q i of the input sequence after passing through the fully connected layer, represents the encoder layer, represents the input sequence embedding vector layer, represents the parameter matrix shared with the input sequence embedding vector matrix, P(Q i ) represents the probability of the output sequence generated by the input sequence Q i after passing through the Softmax layer, and Softmax represents the Softmax layer, f(Qi-1 ) is the (i-1)-th element Q of the input sequence i-1 The vector after passing through the fully connected layer;

[0037] S32 constructs a second model and designs a second network structure layer, including: an input layer, an encoding layer, a fully connected layer, a Softmax layer, and an output layer. The encoding and decoding methods are as follows:

[0038] Construct an encoder based on a bidirectional long short-term network, perform a linear transformation on the last backward hidden state in the encoder, and use its result as the initial hidden state of the decoder. Calculate the current hidden state, and the calculation formula is:

[0039] s t = LSTM(q t-1 , c t-1 , s t-1 )

[0040] Among them, s t represents the hidden state, t is the time step, LSTM is the LSTM network of the encoder, q t-1 is the previous word at time step t, c t-1 is the context vector, s t-1 is the hidden state of the previous time step. The decoder obtains the replication probability of the hidden state between the decoder and the encoder through the attention mechanism, and combines the replication mechanism to complete the prediction of each question word;

[0041] S33 uses the output of the first question model as the input of the second question model to form a sequence question model;

[0042] S34 integrates the Word2Vec model into the input layer of the first question model in the sequence model, clusters the keywords in the input sequence through the Word2Vec model first, and obtains the categories of the keywords;

[0043] S35 constructs the historical data with the information text data as the input sequence, incorporates the categories of the keywords as additional features into the input sequence, and uses the target question as the target output sequence combination as the training set to train the question sequence model to obtain a trained question sequence model.

[0044] Further, the method of embedding the bias analysis function into the question sequence model includes:

[0045] S41 constructs an attention layer after the fully connected layer of the first question model in the question sequence model, and designs a relationship network based on the relationship between the attention mechanism and the receptive field module. The expression of the receptive field attention convolution process is:

[0046] F = Softmax(g i×i(AvgPool(X)))×ReLu(Norm(g k×k (X)))

[0047] Among them, Softmax is the Softmax function, and g i×i is a grouped convolution with a size of i×i, AvgPool is an average pooling operation, X is the output of the fully connected layer, ReLu is an activation function, Norm is a normalization operation, and g k×k (X) represents a gating operation, and k×k is the k channels to which the gating operation is applied to X;

[0048] S42 uses a bias analysis function to calculate weights for each attention head, multiplies the preliminary attention scores of each attention head by their corresponding bias weights, and adjusts their attention scores;

[0049] S43 normalizes each attention head through the Softmax function, uses the normalized attention scores as weights, performs weighted summation on the feature representations of each attention head, and passes the combined feature representation to the next layer of the model.

[0050] Furthermore, the method for calculating the contribution degree of a problem based on the simulation results includes:

[0051]

[0052] Among them, ψ i is the contribution degree of the i-th problem, n is the number of problems, N is the total number of keywords in the full text, df(q i ) is the number of keywords in the i-th problem, is the proportion of the i-th problem based on the total number of keywords in the full text, k i is the sum of the semantic importance degrees of the keywords included in the i-th problem, l i is the length of the i-th problem, is the average length of all problems.

[0053] Furthermore, the method for generating a questionnaire according to the contribution degree includes:

[0054] S51 constructs a threshold division function to divide the problem contribution degrees into levels. The expression of the threshold division function is:

[0055]

[0056] Among them, r is the number of generated problems, is the average value of the problem overlap degree calculated by the word substitution algorithm, is the average length of all problems, Q is the time sensitivity coefficient, is the average contribution degree of all problems, The average number of keywords for all problems, A T is the high-sensitivity coefficient of the problem calculated by the singular value decomposition method;

[0057] Problems with a contribution degree less than f1 are low-contribution problems, problems with a contribution degree in the range [f1, f2] are medium-contribution problems, and problems with a contribution degree greater than f2 are high-contribution problems;

[0058] S53 will use the scheme of retaining high- and low-contribution problems as the questions of the questionnaire and preferentially outputting the problems with higher contribution degrees in the questionnaire as the questionnaire generation scheme to generate the questionnaire.

[0059] The second aspect of the present invention also provides a questionnaire generation system for technical trade measures based on big data, including:

[0060] A data processing module, which acquires the information text data and web traffic data of the trade service platform and performs preprocessing;

[0061] A bias determination module, which performs keyword analysis on the information text data to obtain the semantic importance of keywords, performs interest analysis on the web traffic data to obtain the user interest degree, and sets a bias analysis function based on the keyword semantic importance and the user interest degree;

[0062] A problem generation model construction module, which constructs a first problem generation model and a second problem generation model, concatenates the first problem generation model and the second problem generation model to obtain a problem sequence model, and constructs a problem generation model based on the sequence model according to the focus bias function;

[0063] A questionnaire generation module, which calculates the contribution degree of the problem based on the simulation result, determines the generation scheme in the questionnaire according to the contribution degree of the problem, and completes the generation of the questionnaire based on the questionnaire generation scheme.

[0064] Compared with the prior art, the embodiments of the present invention at least have the following advantages or beneficial effects:

[0065] (1) The present invention provides a method and system for generating a questionnaire for technical trade measures of big data, which combines the analysis of information text data and web traffic data, obtains the semantic importance of keywords and the user bias degree based on data analysis, constructs a bias analysis function to optimize the problem sequence model, and determines the questions and the question generation order in the questionnaire, improving the objectivity and scientificity of the method for generating a questionnaire for technical trade measures.

[0066] (2) The present invention quantifies the contribution degree of the problem, determines the questions in the questionnaire and the question generation order, makes the priority of question generation clearer, improves the efficiency and effect of the questionnaire, and can flexibly adjust the questionnaire content according to different data.

[0067] (3) The present invention combines the semantic importance of keywords and the user interest preference degree, constructs a preference analysis function, and further combines the preference function with the model, enabling the model to pay more attention to information texts with high attention and user interest, being able to capture trade hot issues and reflect the potential preferences of users, and improving the timeliness and market relevance of the generation method of the technical trade measures questionnaire. Description of the Drawings

[0068] Figure 1 It is a flowchart of the steps of the generation method of the technical trade measures questionnaire for the big data of the present invention. Detailed Embodiment

[0069] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0070] Referring to Figure 1 As shown, the first aspect of the present invention provides a generation method of a technical trade measures questionnaire for big data, including:

[0071] S1 Obtain the information text data and web traffic data of the trade service platform, and perform preprocessing;

[0072] In the actual evaluation, a piece of information text data and web traffic data are obtained:

[0073] Among them, the information text data is specifically:

[0074] Since 2023, the new energy vehicle field has faced more than 30 technical trade measures, and these measures involve multiple countries. For example, a certain country issued a notice of G / TBT / N / USA / 1986 on May 2, 2023, updating the advanced clean vehicle II low-emission vehicle and greenhouse gas standards in a certain state, and increasing the requirements for zero-emission vehicles for model years 2027 and later. In addition, a certain country also issued a notice of G / TBT / N / USA / 2034 on August 17, 2023, amending the rules and formulating the Low Emission Vehicle Regulations applicable to light and medium-sized vehicles. This amendment adopted the advanced clean vehicle II standard in a certain state, including the new zero-emission light vehicle sales requirements for model years 2027 - 2032 and later, as well as clarifying changes to the current regulations;

[0075] Among them, the web traffic data includes: http request log frames, click event frames, Web browsing time frames, and scroll depth frames.

[0076] In the actual evaluation, data cleaning and text normalization operations are performed on the obtained information text data and web traffic data.

[0077] S2 Analyze the keywords of the information text data to obtain the semantic importance of the keywords, and analyze the interest degree of the web traffic data to obtain the user interest degree;

[0078] In the actual evaluation, the keywords extracted from the information text data are a certain country, a certain state, advanced clean vehicles, zero-emission light vehicles, and 'G / TBT / N / '. Then, the semantic importance of the keywords is calculated, and the interest degree of the web traffic data is analyzed to obtain the user interest degree.

[0079] S3 Construct a bias analysis function based on the keyword semantic importance and the user interest degree;

[0080] Based on the calculated keyword semantic importance and user interest degree, a bias analysis function is constructed.

[0081] S4 Construct a question sequence model and embed the bias analysis function into the question sequence model;

[0082] In the actual evaluation, design the first network architecture layer and the second network architecture layer to obtain the first question model and the second question model, and design different encoding and decoding methods for the first question model and the second question model. Further, use the question generated by the first question model as the input of the second question model, concatenate and combine them into a sequence model, and obtain the category of the keywords by integrating the Word2Vec model into the input layer of the first question model. Then, construct a training set with the information text data as the input sequence and the target question as the target output sequence to train the question sequence model;

[0083] In the actual evaluation, construct an attention layer after the fully connected layer of the first question model, calculate the weights for each attention head using the bias analysis function, multiply the preliminary attention scores of each attention head by their corresponding bias weights to adjust their attention scores, and perform a normalization operation to weighted sum the feature representations of each attention head and pass the combined feature representation to the next layer.

[0084] S5 Simulate the question sequence model, calculate the contribution degree of the question based on the simulation result, and generate a questionnaire according to the contribution degree.

[0085] In the actual evaluation, simulate the question sequence model to obtain a total of 6 questions, calculate the contribution degree of each question, remove the low-contribution questions through a threshold division function, and retain a total of 4 questions. Generate the questionnaire in descending order of the contribution degree of the questions as follows:

[0086] Do you know the news about the notice G / TBT / N / issued by a certain country?

[0087] What do you think of the relevant measures in the notice G / TBT / N / issued by a certain country?

[0088] What do you think is the long-term impact of the notice G / TBT / N / on advanced clean vehicles?

[0089] What suggestions or opinions do you have for the field of advanced clean vehicles?

[0090] In this embodiment, the method for obtaining the keyword semantic importance by performing keyword analysis on the information text data in step S2 includes:

[0091] S11 Extract all words in the information text data through the TF-IDF algorithm;

[0092] S12 Construct a keyword graph structure for all words. Each node in the graph represents a word, and the edges between nodes represent the semantic relationship between words. The weight of each node is obtained through iterative calculation. The iterative calculation formula is:

[0093]

[0094] Where ω(V i ) is the weight value of node V i , d is the damping coefficient, In(V i ) represents the set of all nodes pointing to node V i , w ji represents the weight value of the edge from node V j to node V i , Out(V j ) represents all nodes pointed to by node V j , w jk represents the weight value of the edge from node V j to node V k , ω(V j ) is the weight value of node j. Sort in descending order according to the node weight value, and take the word corresponding to the node with the highest weight as the final keyword;

[0095] S13 Calculate the semantic importance of the keyword. The calculation formula is:

[0096]

[0097] Where l is the length of keyword m, is the average character length of the keyword, |D| represents the total number of words in the text, N is the number of keywords, Denote the sum of the occurrence frequencies of all keywords in the set {1, 2,... N} except keyword m, C m is the occurrence frequency of keyword m, ω(V m ) represents the node weight value corresponding to keyword m, ‖ω(V m )‖ represents the Euclidean norm of the node weight value, |{j:t i ∈d j}| represents the number of statement sets including keyword m in the text, I m is the inverse document frequency of the statement including keyword m, I k represents the sum of the inverse document frequencies of all keywords in the set {1, 2,... N} except keyword m.

[0098] In this embodiment, the method for obtaining the user interest degree by analyzing the interest degree of the web traffic data includes:

[0099] S21 Obtain the http request log frame and click event frame to extract key data to obtain the user click data of the information, and obtain the page browsing time frame and scroll depth frame to obtain the user information browsing data;

[0100] S22 Introduce a time decay factor, and calculate the user interest preference degree based on the user click data and user information browsing data. The calculation formula is:

[0101]

[0102] λ n =(t s -5)×0.1

[0103] where n is the number of paragraphs of the information text data, is the average click-through rate of the user, is the average browsing time of the user, represents the probability after normalizing the user action vector, μ n is the fatigue factor, representing the reading fatigue value of the user for the reading content as time increases, λ n is the time decay factor, where (t s -5)×0.1 represents a penalty term, indicating that when the browsing time of the user for the information is less than 5s, the user's interest preference degree is reduced.

[0104] In this embodiment, the method for constructing a bias analysis function based on the keyword semantic importance and the user interest degree includes:

[0105] Construct a bias analysis function based on the user interest degree and the semantic importance corresponding to the selected keywords. The expression is:

[0106]

[0107] Among them, k is the number of keywords with semantic importance greater than the average semantic importance, k j and k t are the fuzzy linguistic values of the keyword and the corresponding statement respectively, m is the central scale value of the wavelet function, f is the wavelet function, m ij and σ ij are the translation parameter and dilation parameter of the wavelet function respectively, score j is the semantic importance of the keyword with semantic importance greater than the average semantic importance, score2 is the user interest degree, T is the total number of keywords, is the average node weight value of all keywords, indicates that only the keywords with semantic importance greater than the average semantic importance are analyzed.

[0108] In this embodiment, the method for constructing the problem sequence model includes:

[0109] S31 Construct a first problem model, design a first network structure layer, including: an input layer, an embedding vector layer, an encoding layer, a fully connected layer, a Softmax layer, and an output layer. The processing of data is to input the input sequence as input data into the embedding vector layer, perform encoding processing through the encoding layer, and then complete problem generation through the fully connected layer and the Softmax layer. Among them, the vector output by the encoder with a self-attention mechanism passes through the fully connected layer to obtain the corresponding problem output. The expressions of the encoding and decoding methods are:

[0110] I input =[CLS],Q1,Q2,...,Q m ,[SEP],C1,C2,...,C m ,[SEP]

[0111]

[0112] P(Q i )=Softmax(f(Q i-1 ))

[0113] Among them, I input is the input data, [CLS],Q1,Q2,...,Q m represents the input sequence, m is the number of input sequences, [SEP],C1,C2,...,C n ,[SEP] represents the target problem output sequence, f(Q i ) is the vector of the i-th element Q i of the input sequence after passing through the fully connected layer, Denoted as the encoder layer, Denoted as the input sequence embedding vector layer, Denoted as the parameter matrix shared with the input sequence embedding vector matrix, P(Q i ) Denoted as the input sequence Q i The probability of the output sequence generated after passing through the Softmax layer, Softmax is denoted as the Softmax layer, f(Q i-1 ) is the vector of the (i - 1)-th element of the input sequence Q i-1 After passing through the fully connected layer;

[0114] S32 constructs a second model, designs a second network structure layer, including: an input layer, an encoding layer, a fully connected layer, a Softmax layer, and an output layer. The encoding and decoding methods are as follows:

[0115] Construct an encoder based on a bidirectional long short-term network, perform a linear transformation on the last backward hidden state in the encoder, and use the result as the initial hidden state of the decoder. Calculate the current hidden state, and the calculation formula is:

[0116] s t = LSTM(q t-1 , c t-1 , s t-1 )

[0117] Among them, s t Denoted as the hidden state, t is the time step, LSTM is the LSTM network of the encoder, q t-1 is the previous word at time step t, c t-1 is the context vector, s t-1 is the hidden state of the previous time step. The decoder obtains the copy probability of the hidden state between the decoder and the encoder through the attention mechanism, and combines the copy mechanism to complete the prediction of each question word;

[0118] S33 uses the output of the first question model as the input of the second question model to form a sequence question model;

[0119] S34 integrates the Word2Vec model into the input layer of the first question model in the sequence model, clusters the keywords in the input sequence through the Word2Vec model first, and obtains the categories of the keywords;

[0120] S35 constructs the historical data with the information text data as the input sequence, incorporates the categories of the keywords as additional features into the input sequence, and uses the target question as the target output sequence combination as the training set to train the question sequence model to obtain the trained question sequence model.

[0121] In this embodiment, the method for embedding the bias analysis function into the problem sequence model includes:

[0122] S41 Construct an attention layer after the fully connected layer of the first problem model in the problem sequence model, and design a relationship network based on the attention mechanism and the receptive field module. The expression of the receptive field attention convolution process is:

[0123] F = Softmax(g i×i (AvgPool(X))) × ReLu(Norm(g k×k (X)))

[0124] where Softmax is the Softmax function, g i×i is a grouped convolution of size i×i, AvgPool is the average pooling operation, X is the output of the fully connected layer, ReLu is the activation function, Norm is the normalization operation, and g k×k (X) represents the gating operation, and k×k is the k channels of the gating operation applied to X;

[0125] S42 Use the bias analysis function to calculate the weight for each attention head, multiply the preliminary attention score of each attention head by its corresponding bias weight, and adjust its attention score;

[0126] S43 Normalize each attention head through the Softmax function, use the normalized attention score as the weight, perform weighted summation on the feature representations of each attention head, and pass the combined feature representation to the next layer of the model.

[0127] In this embodiment, the method for calculating the contribution degree of the problem based on the simulation result includes:

[0128]

[0129] where ψ i is the contribution degree of the i-th problem, n is the number of problems, N is the total number of full-text keywords, df(q i ) is the number of keywords in the i-th problem, is the proportion of the i-th problem based on the total number of full-text keywords, k i is the sum of the semantic importance degrees of the keywords included in the i-th problem, l i is the length of the i-th problem, is the average length of all problems.

[0130] In this embodiment, the method for generating the questionnaire according to the contribution degree includes:

[0131] S51 constructs a threshold division function to classify the contribution degrees of problems. The expression of the threshold division function is as follows:

[0132]

[0133] where r is the number of generated problems, is the average value of the problem overlap degree calculated by the word substitution algorithm, is the average length of all problems, Q is the time sensitivity coefficient, is the average contribution degree of all problems, is the average number of keywords of all problems, A T is the high-sensitivity coefficient of the problems calculated by the singular value decomposition method;

[0134] S52 Problems with a contribution degree less than f1 are low-contribution problems, problems with a contribution degree in the range [f1, f2] are medium-contribution problems, and problems with a contribution degree greater than f2 are high-contribution problems. Retain the high- and low-contribution problems as the questions in the questionnaire;

[0135] S53 Use the combination of retaining the high- and low-contribution problems as the questions in the questionnaire and preferentially outputting the problems with higher contribution degrees in the questionnaire as the questionnaire generation plan to generate the questionnaire.

[0136] The second aspect of the present invention also provides a generation system for a technical trade measures questionnaire based on big data, including:

[0137] A data processing module that obtains the information text data and web traffic data of the trade service platform and performs preprocessing;

[0138] A bias determination module that analyzes the keywords of the information text data to obtain the semantic importance of the keywords, analyzes the interest degree of the web traffic data to obtain the user interest degree, and sets a bias analysis function based on the semantic importance of the keywords and the user interest degree;

[0139] A problem generation model construction module that constructs a first problem generation model and a second problem generation model, connects the first problem generation model and the second problem generation model in series to obtain a problem sequence model, and constructs a problem generation model based on the sequence model according to the emphasis bias function;

[0140] A questionnaire generation module that calculates the contribution degree of the problems based on the simulation results, determines the generation plan in the questionnaire according to the contribution degree of the problems, and completes the generation of the questionnaire based on the questionnaire generation plan.

[0141] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of this technology can make various modifications, supplements or use similar methods to replace the specific embodiments described, as long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should all fall within the protection scope of the present invention.

Claims

1. A method for generating a technical trade measures questionnaire based on big data, characterized in that It includes the following steps: S1 Obtain the information text data and web traffic data of the trade service platform, and perform preprocessing; S2 Conduct keyword analysis on the information text data to obtain the semantic importance of keywords, and conduct interest analysis on the web traffic data to obtain user interest; S3 Construct a bias analysis function based on the semantic importance of the keywords and the user interest; S4 Construct a problem sequence model, and embed the bias analysis function into the problem sequence model; S5 Simulate the problem sequence model, calculate the contribution degree of the problem based on the simulation result, and generate a questionnaire according to the contribution degree; The method of constructing the bias analysis function based on the semantic importance of the keywords and the user interest includes: Construct a bias analysis function based on the user interest and the semantic importance corresponding to the selected keywords, and the expression is: Among them, k is the number of keywords whose semantic importance is greater than the average semantic importance, k j and k t are the fuzzy linguistic values of the keyword and the statement corresponding to the keyword respectively, m is the central scale value of the wavelet function, f is the wavelet function, m ij and σ ij are the translation parameter and the dilation parameter of the wavelet function respectively, score j is the semantic importance of the keywords whose semantic importance is greater than the average semantic importance, T is the total number of keywords, score2 is the user interest degree, is the average node weight value of all keywords, indicates that only the keywords whose semantic importance is greater than the average semantic importance are analyzed; The method of constructing the problem sequence model includes: S31 Construct a first problem model, and design a first network structure layer, including: an input layer, an embedding vector layer, an encoding layer, a fully connected layer, a Softmax layer, and an output layer. The processing of the data is to input the input sequence as input data into the embedding vector layer, perform encoding processing through the encoding layer, and then complete problem generation through the fully connected layer and the Softmax layer. Among them, the vector output by the encoder with a self-attention mechanism passes through the fully connected layer to obtain the corresponding problem output. The expressions of the encoding and decoding methods are: I input = [CLS], Q1, Q2,..., Q m , [SEP], C1, C2,..., C m , [SEP] P(Q i ) = Softmax(f(Q i-1 )) Among them, I input is the input data, [CLS], Q1, Q2,..., Q m is expressed as the input sequence, m is the number of input sequences, [SEP], C1, C2,..., C n , [SEP] is expressed as the target problem output sequence, f(Q i ) is the i-th element Q of the input sequence i after passing through the fully connected layer, is expressed as the encoder layer, is expressed as the input sequence embedding vector layer, is expressed as the parameter matrix shared with the input sequence embedding vector matrix, P(Q i ) is expressed as the input sequence Q i is the probability of the output sequence generated after passing through the Softmax layer, Softmax is expressed as the Softmax layer, f(Q i-1 ) is the (i - 1)-th element Q of the input sequence i-1 after passing through the fully connected layer; S32 Construct a second model, and design a second network structure layer, including: an input layer, an encoding layer, a fully connected layer, a Softmax layer, and an output layer. The encoding and decoding methods are: Construct an encoder based on a bidirectional long short-term network, perform a linear transformation on the last backward hidden state in the encoder, and use the result as the initial hidden state of the decoder to calculate the current hidden state. The calculation formula is: s t = LSTM(q t-1 , c t-1 , s t-1 ) Among them, s t is represented as a hidden state, t is the time step, LSTM is the LSTM network of the encoder, q t-1 is the previous word at time step t, c t-1 is the context vector, s t-1 is the hidden state at the previous time step. The decoder obtains the copy probability of the hidden state between the decoder and the encoder through the attention mechanism, and combines the copy mechanism to complete the prediction of each question word; S33 Use the output of the first problem model as the input of the second problem model to form a sequence problem model; S34 Incorporate the Word2Vec model into the input layer of the first problem model in the sequence model, and first cluster the keywords in the input sequence through the Word2Vec model to obtain the categories of the keywords; S35 Construct historical data with the information text data as the input sequence, incorporate the categories of the keywords as additional features into the input sequence, and use the target problem as the target output sequence combination as the training set to train the problem sequence model to obtain a trained problem sequence model.

2. The method for generating a technical trade measure questionnaire for big data according to claim 1, wherein The method of conducting keyword analysis on the information text data to obtain the semantic importance of keywords in step S2 includes: S11 Extract all words in the information text data through the TF-IDF algorithm; S12 Construct a keyword graph structure for all words. Each node in the graph represents a word, and the edges between nodes represent the semantic relationship between words. Calculate the weight of each node through iterative calculation. The iterative calculation formula is: Among them, ω(V i ) is the weight value of node V i , d is the damping coefficient, In(V i ) represents the set of all nodes pointing to node V i , w ji represents the weight value of the edge from node V j to node V i , Out(V j ) represents all nodes pointed to by node V j , w jk represents the weight value of the edge from node V j to node V k , ω(V j ) is the weight value of node j. Sort in descending order according to the node weight value, and use the word corresponding to the node with the highest weight as the final keyword; S13 Calculate the semantic importance of the keywords. The calculation formula is: where l is the length of keyword m, is the average character length of the keyword, |D| represents the total number of words in the text, and N is the number of keywords. represents the sum of the occurrence frequencies of all keywords except keyword m in the set {1, 2,... N}, C m is the occurrence frequency of keyword m, ω(V m ) represents the node weight value corresponding to keyword m, ||ω(V m )|| represents the Euclidean norm of the node weight value, |{j:t i ∈d j}| represents the number of sentence sets including keyword m in the text, I m is the inverse document frequency of the sentence including keyword m, represents the sum of the inverse document frequencies of all keywords except keyword m in the set {1, 2,... N}.

3. The method for generating a technical trade measure questionnaire for big data according to claim 1, wherein The method of conducting interest analysis on the web traffic data to obtain user interest includes: S21 Obtain the HTTP request log frame and click event frame to extract key data to obtain the user click data of the information, and obtain the user information browsing data by obtaining the page viewing time frame and scroll depth frame; S22 Introduce a time decay factor, and calculate the user interest preference degree based on the user click data and user information browsing data. The calculation formula is: λ n = (t s - 5) × 0.1 Among them, n is the number of paragraphs of the information text data, is the average click-through rate of users, is the average browsing time of users, represents the probability after normalizing the user action vector, μ n is the fatigue factor, representing the reading fatigue value of the user for the reading content as time increases, λ n is the time decay factor, where (t s - 5)×0.1 represents the penalty term, indicating that when the browsing time of the user for the information is less than 5s, the user's interest preference is reduced.

4. The method for generating a technical trade measure questionnaire for big data according to claim 1, wherein The method of embedding the bias analysis function into the problem sequence model includes: S41 Construct an attention layer after the fully connected layer of the first problem generation model in the problem sequence model, and design a relationship network based on the attention mechanism and receptive field module. The expression of the receptive field attention convolution process is: F = Softmax(g i×i (AvgPool(X))) × ReLu(Norm(g k×k (X))) Among them, Softmax is the Softmax function, and g i×i is grouped convolution of size i×i, AvgPool is the average pooling operation, X is the output of the fully connected layer, ReLu is the activation function, Norm is the normalization operation, and g k×k (X) represents the gating operation, and k×k are the k channels to which the gating operation is applied to X; S42 Use the bias analysis function to calculate the weight for each attention head, multiply the preliminary attention score of each attention head by its corresponding bias weight, and adjust its attention score; S43 Normalize each attention head through the Softmax function, use the normalized attention score as the weight, perform weighted summation on the feature representations of each attention head, and pass the combined feature representation to the next layer of the model.

5. The method for generating a technical trade measures questionnaire for big data according to claim 1, characterized in that The method of calculating the contribution degree of the problem based on the simulation result includes: Among them, ψ i is the contribution degree of the i-th question, n is the number of questions, N is the total number of keywords in the full text, df(q i ) is the number of keywords in the i-th question, is the proportion of the keywords of the i-th question based on the total number of keywords in the full text, k i is the sum of the semantic importance degrees of the keywords included in the i-th question, l i is the length of the i-th question, is the average length of all questions.

6. The method for generating a technical trade measure questionnaire for big data according to claim 1, characterized in that The method of generating a questionnaire according to the contribution degree includes: S51 Construct a threshold division function to divide the problem contribution degree into levels. The expression of the threshold division function is: where r is the number of generated questions, is the average value of the question overlap calculated by the word substitution algorithm, is the average length of all questions, Q is the time sensitivity coefficient, is the average contribution of all questions, is the average number of keywords of all questions, A T is the high-sensitivity coefficient of the questions calculated by the singular value decomposition method; S52 Problems with a contribution degree less than f1 are low contribution degree problems, problems with a contribution degree in the range [f1, f2] are medium contribution degree problems, and problems with a contribution degree greater than f2 are high contribution degree problems; S53 Use the scheme of retaining high and low contribution degree problems as the questions of the questionnaire and giving priority to outputting the problems with higher contribution degrees in the questionnaire as the questionnaire generation scheme to generate the questionnaire.

7. A generation system for a technical trade measures questionnaire based on big data, which is used to execute the method described in any one of claims 1-6, characterized in that The system includes: A data processing module that obtains the information text data and web traffic data of the trade service platform and performs preprocessing; A bias determination module that performs keyword analysis on the information text data to obtain the keyword semantic importance, performs interest analysis on the web traffic data to obtain the user interest degree, and sets a bias analysis function based on the keyword semantic importance and user interest degree; A module for constructing a problem generation model that constructs a first problem generation model and a second problem generation model, connects the first problem generation model and the second problem generation model in series to obtain a problem sequence model, and constructs a problem generation model based on the sequence model according to the emphasis bias function; A questionnaire generation module that calculates the contribution degree of the problem based on the simulation result, determines the generation scheme in the questionnaire according to the contribution degree of the problem, and completes the generation of the questionnaire based on the questionnaire generation scheme.

Citation Information

Patent Citations

  • Personalized text sequencing and recommending method for network users

    CN104298732A

  • Method for generating technical trade measure questionnaire

    CN115906792A