Content management method and system based on big data

By using deep learning technology to perform word-grained semantic coding, semantic explicit modeling and emotion recognition in the content management method, the problems of insufficient deep semantic understanding of text and neglecting emotional factors in the existing technology are solved, and more accurate and personalized input prompt word generation is achieved.

CN120045694APending Publication Date: 2025-05-27SHANDONG YOUTH UNIV OF POLITICAL SCI +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510182135.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-24
Filing Date
2025-02-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing content management method based on prompt word Q&A interaction lacks an understanding of the deep context semantics of the text when inputting prompt word extraction. Especially when dealing with vague or vague queries, its accuracy and response efficiency may be affected, and the impact of emotional factors on input prompt word extraction is not considered.

Method used

The natural language processing technology based on deep learning is adopted to perform word-grained semantic coding on input problems, semantic explicit modeling and global semantic information aggregation, and emotional recognition is performed based on the global semantic features of the input problems, and input prompt words are intelligently generated based on the semantics and emotional features.

Benefits of technology

Effectively capture semantic details and emotional tendencies in input problems, improve the accuracy and relevance of input prompt words, be able to handle clear and vague queries, and provide personalized and precise content management services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045694A_ABST
    Figure CN120045694A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of content management, and particularly discloses a content management method and system based on big data, and the method comprises the steps: carrying out the word granularity semantic coding of an input question through employing a natural language processing technology based on deep learning, so as to obtain the semantic feature representation of each word unit; according to the method, input questions are extracted, semantic explicit modeling and global semantic information aggregation are carried out on the input questions, emotion recognition is carried out based on global semantic features of the input questions, feature coding representation of emotion type tags is obtained, and then input prompt words of the input questions of the user are intelligently generated by combining the global semantic features and emotion type features of the input questions. Therefore, semantic details and emotional tendencies in the input questions can be effectively captured, so that the accuracy and correlation of the input cue words are improved, the problems which are implicit or have emotional colors are effectively understood and responded, and more personalized and accurate content management services are provided for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of content management, and more specifically, to a content management method and system based on big data. Background Art

[0002] With the rapid development of Internet technology and the explosive growth of information volume, content management has become one of the important challenges faced by enterprises and individuals. Especially in the big data environment, how to effectively manage and utilize massive data resources and improve the availability and value of data has become the focus of research.

[0003] In response to this, the invention patent with the publication number CN116860935A proposes a content management method based on prompt word question-and-answer interaction. By obtaining the user input question on the interaction interface, extracting common words and filtering stop words from the input question to obtain input prompt words, and then successively performing keyword verification, template expectation verification, personalized verification, and emotional anomaly judgment based on the input prompt words, and selectively generating corresponding answers, so that users can quickly select the most suitable module for data display and function configuration, greatly improving the configuration and release efficiency of operators for digital content management.

[0004] However, when extracting input prompt words in the above solution, only by the methods of common word extraction and stop word filtering, it mainly relies on the surface structure of the text to parse the user's intention, lacking the ability to understand the deep context semantics of the text. Especially when dealing with fuzzy or ambiguous queries, its accuracy and response efficiency may be affected. In addition, the emotional tendency in the user input question is also one of the important factors affecting the extraction of input prompt words. However, the above solution does not consider the influence of emotional factors on the extraction of input prompt words. Therefore, when dealing with questions involving emotional colors, it may not be able to accurately capture the actual needs of users, thereby reducing the user experience.

[0005] Therefore, an optimized content management method and system based on big data are expected. Summary of the Invention

[0006] To solve the above technical problems, the present application is proposed. An embodiment of the present application provides a content management method and system based on big data, which uses natural language processing technology based on deep learning to perform word-level semantic encoding on the input question to obtain semantic feature representations of each word unit. Then, semantic explicit modeling and global semantic information aggregation are performed on it, and sentiment recognition is performed based on the global semantic features of the input question to obtain a feature encoding representation of the sentiment type label. Furthermore, combining the global semantic features and sentiment type features of the input question, input prompt words for the user input question are intelligently generated. In this way, semantic details and sentiment tendencies in the input question can be effectively captured, thereby improving the accuracy and relevance of the input prompt words. It can not only process explicit query inputs, but also effectively understand and respond to ambiguous or emotionally charged questions, thus providing users with more personalized and accurate content management services.

[0007] Correspondingly, according to one aspect of the present application, there is provided a content management method based on big data, which includes: obtaining an input question of an interaction interface, extracting a prompt word for the input question to obtain an input prompt word for the input question; obtaining an interaction database of the interaction interface, and using the interaction database to determine whether the input prompt word passes keyword verification; when the input prompt word passes the keyword verification, using the interaction database to generate a feedback answer for the input prompt word, wherein extracting a prompt word for the input question to obtain an input prompt word for the input question includes: performing word-level semantic encoding on the input question to obtain a sequence of input question word-level semantic encoding vectors; performing semantic explicit modeling on the sequence of input question word-level semantic encoding vectors to obtain a sequence of input question word-level semantic enhanced encoding vectors; performing sentiment recognition on the sequence of input question word-level semantic enhanced encoding vectors to obtain a sentiment type label encoding vector; generating a prompt word based on the sentiment type label encoding vector and the sequence of input question word-level semantic enhanced encoding vectors to obtain an input prompt word for the input question.

[0008] In the above content management method based on big data, performing semantic explicit modeling on the sequence of input question word-level semantic encoding vectors includes: calculating semantic saliency description factors of each input question word-level semantic encoding vector in the sequence of input question word-level semantic encoding vectors to obtain a sequence of input question word-level semantic saliency description factors; based on the sequence of input question word-level semantic saliency description factors, performing hierarchical mask modulation on the sequence of input question word-level semantic encoding vectors to obtain the sequence of input question word-level semantic enhanced encoding vectors.

[0009] In the above content management method based on big data, performing word-level semantic encoding on the input question to obtain a sequence of input question word-level semantic encoding vectors includes: performing word segmentation on the input question and then using a semantic encoder including a Bert model to obtain the sequence of input question word-level semantic encoding vectors.

[0010] In the above content management method based on big data, calculating semantic significance description factors for each input question word-level semantic encoding vector in the sequence of input question word-level semantic encoding vectors to obtain a sequence of input question word-level semantic significance description factors includes: mapping each input question word-level semantic encoding vector in the sequence of input question word-level semantic encoding vectors to a hyperbolic space to obtain a sequence of mapped input question word-level semantic encoding vectors; calculating the squared Poincaré norm of each mapped input question word-level semantic encoding vector in the sequence of mapped input question word-level semantic encoding vectors as the semantic significance description factor to obtain the sequence of input question word-level semantic significance description factors.

[0011] In the above content management method based on big data, based on the sequence of input question word-level semantic significance description factors, performing hierarchical mask modulation on the sequence of input question word-level semantic encoding vectors to obtain a sequence of input question word-level semantic enhanced encoding vectors includes: performing hierarchical mask modulation on the sequence of input question word-level semantic significance description factors to obtain a sequence of input question word-level semantic significance modulation weights; using each input question word-level semantic significance modulation weight in the sequence of input question word-level semantic significance modulation weights as a weight to perform weighted modulation on the corresponding input question word-level semantic encoding vector in the sequence of input question word-level semantic encoding vectors respectively to obtain the sequence of input question word-level semantic enhanced encoding vectors.

[0012] In the above content management method based on big data, performing hierarchical mask modulation on the sequence of input question word-level semantic significance description factors to obtain a sequence of input question word-level semantic significance modulation weights includes: inputting the sequence of input question word-level semantic significance description factors into a normalization process based on the softmax function to obtain a sequence of normalized input question word-level semantic significance description factors; inputting the sequence of normalized input question word-level semantic significance description factors into a multi-level mask function to obtain the sequence of input question word-level semantic significance modulation weights.

[0013] In the above-mentioned big data-based content management method, inputting the sequence of the normalized input problem word granularity semantic saliency description factors into a multi-level masking function to obtain the sequence of the input problem word granularity semantic saliency modulation weights includes: setting a first masking threshold and a second masking threshold, where the second masking threshold is twice the first masking threshold; comparing each of the normalized input problem word granularity semantic saliency description factors in the sequence of the normalized input problem word granularity semantic saliency description factors with the first masking threshold and the second masking threshold. If the normalized input problem word granularity semantic saliency description factor is greater than the second masking threshold, it is weighted and amplified; if the normalized input problem word granularity semantic saliency description factor is greater than the first masking threshold and less than or equal to the second masking threshold, it remains unchanged; if the normalized input problem word granularity semantic saliency description factor is less than or equal to the first masking threshold, it is weighted and suppressed, so as to obtain the sequence of the input problem word granularity semantic saliency modulation weights.

[0014] In the above-mentioned big data-based content management method, performing sentiment recognition on the sequence of the input problem word granularity semantic reinforcement encoding vectors to obtain a sentiment type label encoding vector includes: concatenating the sequence of the input problem word granularity semantic reinforcement encoding vectors to obtain an input problem global semantic encoding vector; inputting the input problem global semantic encoding vector into a sentiment recognition module based on a classifier to obtain the sentiment type label encoding vector.

[0015] In the above-mentioned big data-based content management method, generating a prompt word for the input problem based on the sentiment type label encoding vector and the sequence of the input problem word granularity semantic reinforcement encoding vectors to obtain the input prompt word for the input problem includes: concatenating the input problem global semantic encoding vector and the sentiment type label encoding vector to obtain an input problem semantic-sentiment concatenated encoding vector; inputting the input problem semantic-sentiment concatenated encoding vector into a prompt word generator based on a decoder to obtain the input prompt word for the input problem.

[0016] According to another aspect of the present application, there is provided a big data-based content management system, which includes: a word-level semantic encoding module for performing word-level semantic encoding on an input question to obtain a sequence of input question word-level semantic encoding vectors; a semantic explicit modeling module for performing semantic explicit modeling on the sequence of input question word-level semantic encoding vectors to obtain a sequence of input question word-level enhanced semantic encoding vectors; an input question sentiment recognition module for performing sentiment recognition on the sequence of input question word-level enhanced semantic encoding vectors to obtain a sentiment type label encoding vector; and a prompt word generation module for generating a prompt word for the input question based on the sentiment type label encoding vector and the sequence of input question word-level enhanced semantic encoding vectors to obtain an input prompt word for the input question.

[0017] Compared with the prior art, the big data-based content management method and system provided by the present application use natural language processing technology based on deep learning to perform word-level semantic encoding on an input question to obtain semantic feature representations of each word unit. Then, semantic explicit modeling and global semantic information aggregation are performed on it, and sentiment recognition is performed based on the global semantic features of the input question to obtain a feature encoding representation of the sentiment type label. Furthermore, by combining the global semantic features and sentiment type features of the input question, an input prompt word for the user input question is intelligently generated. In this way, semantic details and sentiment tendencies in the input question can be effectively captured, thereby improving the accuracy and relevance of the input prompt word. It can not only handle explicit query inputs but also effectively understand and respond to vague or emotionally charged questions, thus providing users with more personalized and accurate content management services. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. They are used together with the embodiments of the present application to explain the present application and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0019] Figure 1 It is a flowchart of a big data-based content management method according to an embodiment of the present application.

[0020] Figure 2 It is a schematic diagram of data flow of a big data-based content management method according to an embodiment of the present application.

[0021] Figure 3 It is a flowchart of step S120 in a big data-based content management method according to an embodiment of the present application.

[0022] Figure 4It is a flowchart of step S130 in the big data-based content management method according to an embodiment of the present application.

[0023] Figure 5 It is a flowchart of step S140 in the big data-based content management method according to an embodiment of the present application.

[0024] Figure 6 It is a block diagram of a big data-based content management system according to an embodiment of the present application. Detailed implementation manners

[0025] Next, the embodiments of the present application will be described in more detail with reference to the accompanying drawings, and the above and other objects, features, and advantages of the present application will become more obvious. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein.

[0026] As mentioned in the above background art, patent CN116860935A proposes a content management method based on prompt-word question-and-answer interaction, which includes: obtaining an input question on an interaction interface, extracting prompt words from the input question to obtain input prompt words of the input question; obtaining an interaction database of the interaction interface, and using the interaction database to determine whether the input prompt words pass keyword verification; when the input prompt words pass the keyword verification, using the interaction database to generate a feedback answer for the input prompt words.

[0027] The above content management method based on prompt-word question-and-answer interaction extracts input prompt words of the user's input question, and sequentially performs keyword verification, template expectation verification, personalized verification, and emotion anomaly judgment based on the input prompt words, and selectively generates corresponding answers, enabling users to quickly select the most suitable module for data display and function configuration, and greatly improving the configuration and release efficiency of operators for digital content management.

[0028] However, when extracting input prompt words in the above solution, only the methods of common word extraction and stop word filtering are used, lacking the ability to understand the deep context semantics of the text, and its accuracy and response efficiency may be affected. In addition, the emotional tendency in the user's input question is also one of the important factors affecting the extraction of input prompt words. And the above solution does not consider the influence of emotional factors on the extraction of input prompt words. Therefore, when dealing with questions involving emotional colors, it may not be able to accurately capture the actual needs of users, thereby reducing the user experience.

[0029] In view of the above technical problems, the present application proposes an optimized big data-based content management method. It uses natural language processing technology based on deep learning to perform word-level semantic encoding on the input question to obtain semantic feature representations of each word unit. Then, it performs semantic explicit modeling and global semantic information aggregation on them, and conducts sentiment recognition based on the global semantic features of the input question to obtain a feature encoding representation of the sentiment type label. Furthermore, by combining the global semantic features and sentiment type features of the input question, it intelligently generates an input prompt word for the user's input question. In this way, semantic details and sentiment tendencies in the input question can be effectively captured, thereby improving the accuracy and relevance of the input prompt word. It can not only handle explicit query inputs but also effectively understand and respond to vague or sentiment-laden questions, thus providing users with more personalized and accurate content management services.

[0030] Figure 1 FIG. is a flowchart of a big data-based content management method according to an embodiment of the present application. Figure 2 FIG. is a schematic diagram of data flow of a big data-based content management method according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the big data-based content management method according to an embodiment of the present application includes the steps of: S110, performing word-level semantic encoding on the input question to obtain a sequence of input question word-level semantic encoding vectors; S120, performing semantic explicit modeling on the sequence of input question word-level semantic encoding vectors to obtain a sequence of input question word-level semantic enhanced encoding vectors; S130, performing sentiment recognition on the sequence of input question word-level semantic enhanced encoding vectors to obtain a sentiment type label encoding vector; S140, generating a prompt word based on the sentiment type label encoding vector and the sequence of input question word-level semantic enhanced encoding vectors to obtain an input prompt word for the input question.

[0031] In the above content management method based on big data, in step S110, the input question is subjected to word-level semantic encoding to obtain a sequence of input question word-level semantic encoding vectors. In a specific example of the present application, step S110 includes: performing word segmentation on the input question and then using a semantic encoder including a Bert model to obtain the sequence of input question word-level semantic encoding vectors. That is, in order to achieve a fine-grained understanding of the deep semantic information of the input question, first, the input question is subjected to word segmentation to refine the granularity of semantic understanding, and then a semantic encoder including a Bert model is used to process each word unit to extract the semantic feature representation of each word unit. Those skilled in the art should know that the Bert model uses a bidirectional Transformer architecture to understand the context of the text, can effectively capture the complex relationships and deep semantics between words, and more accurately understand the specific meanings of each word unit in the context of the input question, so as to obtain a sequence of input question word-level semantic encoding vectors, providing a more reliable data basis for subsequent prompt generation.

[0032] In the above content management method based on big data, in step S120, semantic explicit modeling is performed on the sequence of input question word-level semantic encoding vectors to obtain a sequence of input question word-level semantic enhanced encoding vectors. It should be understood that considering that each word unit in the input question usually has different semantic importance, therefore, in order to better understand the overall semantic structure of the input question, the present application further performs semantic explicit modeling on the sequence of input question word-level semantic encoding vectors to highlight key word information and suppress the interference of unimportant information, thereby improving the feature discrimination degree and further enhancing the semantic expression ability of the input question.

[0033] Figure 3 The flowchart of step S120 in the content management method based on big data according to an embodiment of the present application is as follows. As Figure 3 shown, step S120 includes: S121, calculating the semantic saliency description factors of each input question word-level semantic encoding vector in the sequence of input question word-level semantic encoding vectors to obtain a sequence of input question word-level semantic saliency description factors; S122, based on the sequence of input question word-level semantic saliency description factors, performing hierarchical mask modulation on the sequence of input question word-level semantic encoding vectors to obtain the sequence of input question word-level semantic enhanced encoding vectors.

[0034] Specifically, the step S121 includes: mapping each input problem word granular semantic encoding vector in the sequence of input problem word granular semantic encoding vectors to the hyperbolic space to obtain a sequence of mapped input problem word granular semantic encoding vectors; calculating the squared Poincaré norm of each mapped input problem word granular semantic encoding vector in the sequence of mapped input problem word granular semantic encoding vectors as the semantic salience description factor to obtain a sequence of input problem word granular semantic salience description factors.

[0035] That is, considering that compared with the Euclidean space, the hyperbolic space can more effectively embed more hierarchical structure information, which helps to better understand the semantic hierarchical structure of the input problem. Therefore, in this application, each input problem word granular semantic encoding vector is first converted from the Euclidean space to the hyperbolic space. In the hyperbolic space, as the distance increases, the space volume grows exponentially. Based on this property, the semantic hierarchical structure of the input problem can be better maintained. Then, in the hyperbolic space, the squared Poincaré norm of each input problem word granular semantic encoding vector is further calculated to measure the importance or contribution degree of the feature, so as to obtain a sequence of input problem word granular semantic salience description factors.

[0036] Specifically, the step S122 includes: performing hierarchical mask modulation on the sequence of input problem word granular semantic salience description factors to obtain a sequence of input problem word granular semantic salience modulation weights; using each input problem word granular semantic salience modulation weight in the sequence of input problem word granular semantic salience modulation weights as a weight to perform weighted modulation on the corresponding input problem word granular semantic encoding vector in the sequence of input problem word granular semantic encoding vectors respectively to obtain a sequence of input problem word granular semantic enhanced encoding vectors.

[0037] In a specific example of this application, performing hierarchical mask modulation on the sequence of input problem word granular semantic salience description factors to obtain a sequence of input problem word granular semantic salience modulation weights includes: inputting the sequence of input problem word granular semantic salience description factors into a normalization process based on the softmax function to obtain a sequence of normalized input problem word granular semantic salience description factors; inputting the sequence of normalized input problem word granular semantic salience description factors into a multi-level mask function to obtain the sequence of input problem word granular semantic salience modulation weights.

[0038] In a specific example of the present application, inputting the sequence of the normalized input question word granularity semantic salience description factors into a multi-level masking function to obtain the sequence of the input question word granularity semantic salience modulation weights includes: setting a first masking threshold and a second masking threshold, where the second masking threshold is twice the first masking threshold; comparing each normalized input question word granularity semantic salience description factor in the sequence of the normalized input question word granularity semantic salience description factors with the first masking threshold and the second masking threshold. If the normalized input question word granularity semantic salience description factor is greater than the second masking threshold, it is weighted and amplified; if the normalized input question word granularity semantic salience description factor is greater than the first masking threshold and less than or equal to the second masking threshold, it remains unchanged; if the normalized input question word granularity semantic salience description factor is less than or equal to the first masking threshold, it is weighted and suppressed, thereby obtaining the sequence of the input question word granularity semantic salience modulation weights.

[0039] That is, through three-level hierarchical masking modulation of the sequence of the input question word granularity semantic salience description factors, semantic salience description factors at different levels undergo feature scaling based on different scales, so as to, while retaining different word granularity feature components, pay different degrees of attention to their features, thereby better highlighting important features and suppressing irrelevant features. Finally, the obtained input question word granularity semantic salience modulation weights are used to perform weighted modulation on the original input question word granularity semantic encoding vectors, so as to emphasize the semantic feature representations closely related to the input question theme, thereby obtaining a sequence of more representative and discriminative input question word granularity semantic enhanced encoding vectors.

[0040] Correspondingly, the step S120 includes: processing the sequence of the input question word granularity semantic encoding vectors with the following feature enhancement formula to obtain the sequence of the input question word granularity semantic enhanced encoding vectors, where the feature enhancement formula is:

[0041] X = {x 1 , x 2 ,..., x i ,..., x n}

[0042] h i = W 1 x i W 2

[0043]

[0044] w i = mask(o i ′)

[0045]

[0046] Y = {y 1 , y 2 ,..., y i ,..., y n}

[0047] y i = w i ·x i

[0048] where X represents the sequence of the input problem word granular semantic encoding vectors, x 1 , x 2 , x i and x n respectively represent the first, second, i-th, and n-th input problem word granular semantic encoding vectors in the sequence of the input problem word granular semantic encoding vectors, n is the number of the input problem word granular semantic encoding vectors, W 1 and W 2 are different weight matrices for projecting the input problem word granular semantic encoding vectors from the Euclidean space to the hyperbolic space, h i is the i-th mapped input problem word granular semantic encoding vector, ‖·‖ 2 represents the square of the norm of the feature vector, log represents the logarithmic function with base 2, exp(·) represents the exponential function with base e, o i represents the i-th input problem word granular semantic salience description factor, o i ′ represents the i-th normalized input problem word granular semantic salience description factor, θ represents the first mask threshold, λ and η are different weight parameters, and λ > 1, η < 1, mask(·) is a multi-level mask function, w i is the i-th input problem word granular semantic salience modulation weight, Y is the sequence of the input problem word granular semantic enhanced encoding vectors, y 1 , y 2 , y i and y n respectively represent the first, second, i-th, and n-th input problem word granular semantic enhanced encoding vectors in the sequence of the input problem word granular semantic enhanced encoding vectors.

[0049] In the above content management method based on big data, in step S130, sentiment recognition is performed on the sequence of the input problem word granular semantic enhanced encoding vectors to obtain a sentiment type label encoding vector. Wherein, Figure 4It is a flowchart of step S130 in the content management method based on big data according to an embodiment of the present application. As Figure 4 shown, the step S130 includes: S131, concatenating the sequence of the input question word granularity semantic enhancement encoding vectors to obtain an input question global semantic encoding vector; S132, inputting the input question global semantic encoding vector into a sentiment recognition module based on a classifier to obtain the sentiment type label encoding vector.

[0050] Specifically, in the step S131, the sequence of the input question word granularity semantic enhancement encoding vectors is concatenated to obtain an input question global semantic encoding vector. That is, in order to achieve a global understanding of the input question, the present application further performs feature fusion on the sequence of the input question word granularity semantic enhancement encoding vectors through feature concatenation, forming a global input question global semantic encoding vector, so as to realize the integration of the global context semantic information of the input question, and provide a comprehensive semantic background guidance for subsequent sentiment recognition and input prompt word generation.

[0051] Specifically, in the step S132, the input question global semantic encoding vector is input into a sentiment recognition module based on a classifier to obtain the sentiment type label encoding vector. It should be understood that considering that the sentiment tendency in the user input question is also one of the important factors affecting the extraction of input prompt words. Therefore, the present application further introduces the sentiment information of the input question to improve the accuracy of prompt word generation. Specifically, the present application uses a classifier to construct a sentiment recognition module to perform sentiment recognition on the input question global semantic encoding vector. In the embodiment of the present application, the classifier is based on a multi-layer perceptron (MLP) structure, and through multi-layer feature learning on the input question global semantic encoding vector, it can effectively capture the emotional color when the user asks a question, identify the sentiment type in the question, such as positive, negative or neutral, etc., and then output the feature encoding representation of the corresponding sentiment type label, that is, the sentiment type label encoding vector.

[0052] Specifically, in the technical solution of the present application, each input problem word granular semantic encoding vector in the sequence of the input problem word granular semantic encoding vectors respectively represents the text semantic embedding encoding features of each word determined by word segmentation processing of the input problem. When enhancing the feature sequence based on explicit modeling of feature descriptions, although it can improve the text semantic salience of each input problem word granular semantic encoding vector in the sequence of the input problem word granular semantic encoding vectors, it will also cause the overall semantic feature manifold of the input problem global semantic encoding vector obtained by concatenating the sequence of the input problem word granular semantic enhanced encoding vectors to have a sharpened feature manifold framework in the high-dimensional feature space due to expression sparsity and distribution imbalance, thereby affecting the accuracy of the sentiment type label encoding vector obtained through the sentiment recognition module based on the classifier.

[0053] Based on this, in a preferred embodiment, during the process of inputting the input problem global semantic encoding vector into the sentiment recognition module based on the classifier to obtain the sentiment type label encoding vector, feature manifold framework modulation is performed on the input problem global semantic encoding vector, which includes:

[0054] Using a random permutation operator to perform unordered reconstruction on the input problem global semantic encoding vector to obtain an input problem global semantic feature random reconstruction encoding vector, denoted as where ρ(v c ) is the random permutation operator, v c represents the input problem global semantic encoding vector, represents matrix multiplication, and v s represents the input problem global semantic feature random reconstruction encoding vector;

[0055] Calculating the outer product of the input problem global semantic encoding vector to obtain a Banach space field encoding matrix, denoted as: where v c T represents the transposed vector of the input problem global semantic encoding vector, and B represents the Banach space field encoding matrix;

[0056] Multiplying the input problem global semantic feature random reconstruction encoding vector by the Banach space field encoding matrix and mapping it to the Banach space field to obtain an input problem global semantic feature numerical relationship heterogeneous encoding vector, denoted as: where C represents the input problem global semantic feature numerical relationship heterogeneous encoding vector;

[0057] Performing morphological discretization on the input problem global semantic feature numerical relationship heterogeneous encoding vector and the transposed vector of the input problem global semantic encoding vector to obtain an input problem global semantic feature asymmetric topological matrix, denoted as Among them, SVD k represents the truncated decomposition that retains the first k singular values, and M represents the asymmetric topological matrix of the global semantic features of the input problem;

[0058] Multiply the global semantic encoding vector of the input problem by the asymmetric topological matrix of the global semantic features of the input problem to obtain an optimized global semantic encoding vector of the input problem, which is expressed as: where v c ′ represents the optimized global semantic encoding vector of the input problem. Finally, input the optimized global semantic encoding vector of the input problem into the sentiment recognition module based on the classifier to obtain the sentiment type label encoding vector.

[0059] Here, by using the random permutation operator to reconstruct the global semantic encoding vector of the input problem in an unordered manner and map it into the Banach space defined by the outer product, a meaningful metric of the numerical relationship of the feature set within the asymmetric topology can be realized. Based on this, a feature space with an asymmetric topology is constructed by encoding the absolute coordinates of the feature vector, and the morphological discretization of the low-dimensional surface of the feature vector within the feature space is performed based on vector queries. In this way, the probability of representing the feature vector due to the sharpening framework can be avoided, thereby increasing the accuracy of the sentiment type label encoding vector obtained by inputting it into the sentiment recognition module based on the classifier.

[0060] In the above content management method based on big data, in step S140, prompt words are generated based on the sequence of the sentiment type label encoding vector and the input problem word granularity semantic enhancement encoding vector to obtain the input prompt words of the input problem. Among them, Figure 5 is the flowchart of step S140 in the content management method based on big data according to an embodiment of the present application. As Figure 5 shown, step S140 includes: S141, concatenating the global semantic encoding vector of the input problem and the sentiment type label encoding vector to obtain an input problem semantic-sentiment concatenated encoding vector; S142, inputting the input problem semantic-sentiment concatenated encoding vector into the prompt word generator based on the decoder to obtain the input prompt words of the input problem.

[0061] Specifically, in step S141, the global semantic encoding vector of the input problem and the sentiment type label encoding vector are concatenated to obtain an input problem semantic-sentiment concatenated encoding vector. That is, the global semantic encoding vector of the input problem and the sentiment type label encoding vector are also fused with information in the form of feature concatenation, so that when generating input prompt words, not only the semantic content of the problem is considered, but also the user's emotional state can be taken into account, thereby more accurately understanding the user's actual needs.

[0062] Specifically, in step S142, the input problem semantic-emotional cascade encoding vector is input into a decoder-based prompt word generator to obtain the input prompt word of the input problem. In the technical solution of this application, the prompt word generator adopts a sequence-to-sequence (Seq2Seq) model. Among them, the decoder part decodes the input problem semantic-emotional cascade encoding vector word by word to intelligently generate an input prompt word that is highly relevant to the user's input problem and can reflect the emotional tendency. In this way, not only can accurate keyword prompts be provided, but also when the user asks a vague or emotional question, a more appropriate and user-friendly response can be given.

[0063] In summary, the big data-based content management method according to the embodiments of this application is described. It uses deep learning-based natural language processing technology to perform word-level semantic encoding on the input problem to obtain the semantic feature representations of each word unit. Then, semantic explicit modeling and global semantic information aggregation are performed on it, and emotion recognition is performed based on the global semantic features of the input problem to obtain the feature encoding representation of the emotion type label. Furthermore, by combining the global semantic features and emotion type features of the input problem, the input prompt word of the user's input problem is intelligently generated. In this way, the semantic details and emotional tendencies in the input problem can be effectively captured, thereby improving the accuracy and relevance of the input prompt word. It can not only process clear query inputs, but also effectively understand and respond to vague or emotionally colored questions, thus providing a more personalized and accurate content management service for users.

[0064] Figure 6 The block diagram of the big data-based content management system according to the embodiments of this application is shown as Figure 6 As shown, the big data-based content management system 100 according to the embodiments of this application includes: a word-level semantic encoding module 110, configured to perform word-level semantic encoding on the input problem to obtain a sequence of input problem word-level semantic encoding vectors; a semantic explicit modeling module 120, configured to perform semantic explicit modeling on the sequence of input problem word-level semantic encoding vectors to obtain a sequence of input problem word-level enhanced encoding vectors; an input problem emotion recognition module 130, configured to perform emotion recognition on the sequence of input problem word-level enhanced encoding vectors to obtain an emotion type label encoding vector; and a prompt word generation module 140, configured to generate a prompt word based on the emotion type label encoding vector and the sequence of input problem word-level enhanced encoding vectors to obtain the input prompt word of the input problem.

[0065] Here, those skilled in the art can understand that the specific operations of each module in the above big data-based content management system have been described above with reference to Figures 1 to 5The description of the content management method based on big data has been introduced in detail, and therefore, its repeated description will be omitted.

[0066] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and easy understanding, rather than limitations. The above details do not limit the present invention to necessarily adopt the above specific details for implementation.

[0067] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the module division is only a logical function division, and there may be other division methods in actual implementation. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0068] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.

[0069] In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.

[0070] Finally, it should be noted that the above description has been given for the purpose of illustration and description. In addition, the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A content management method based on big data, comprising: Acquire an input question of the interactive interface, extract prompt words for the input question, and obtain an input prompt word for the input question; Acquire an interactive database of the interactive interface, and use the interactive database to determine whether the input prompt word passes keyword verification; when the input prompt word passes the keyword verification, use the interactive database to generate a feedback answer to the input prompt word, characterized in that extracting prompt words from the input question to obtain the input prompt word of the input question includes: Performing word-granular semantic encoding on the input question to obtain a sequence of word-granular semantic encoding vectors of the input question; Performing semantic explicit modeling on the sequence of input question word granular semantic encoding vectors to obtain a sequence of input question word granular semantic enhanced encoding vectors; Performing sentiment recognition on the sequence of the input question word granular semantic enhancement encoding vectors to obtain a sentiment type label encoding vector; Generating prompt words based on the sequence of the emotion type label encoding vector and the input question word granularity semantic enhancement encoding vector to obtain an input prompt word for the input question; Among them, semantic explicit modeling is performed on the sequence of input question word granularity semantic coding vectors, including: calculating the semantic significance description factor of each input question word granularity semantic coding vector in the sequence of input question word granularity semantic coding vectors to obtain a sequence of input question word granularity semantic significance description factors; based on the sequence of input question word granularity semantic significance description factors, hierarchical mask modulation is performed on the sequence of input question word granularity semantic coding vectors to obtain a sequence of input question word granularity semantic reinforcement coding vectors.

2. The content management method based on big data according to claim 1, characterized in that: Performing word-granular semantic encoding on the input question to obtain a sequence of word-granular semantic encoding vectors of the input question includes: After word segmentation, the input question is passed through a semantic encoder including a Bert model to obtain a sequence of word-granular semantic encoding vectors of the input question.

3. The content management method based on big data according to claim 2 is characterized in that: Calculating the semantic significance description factor of each input question word granularity semantic encoding vector in the sequence of input question word granularity semantic encoding vectors to obtain a sequence of input question word granularity semantic significance description factors, including: Mapping each input question word granular semantic encoding vector in the sequence of input question word granular semantic encoding vectors to a hyperbolic space to obtain a sequence of mapped input question word granular semantic encoding vectors; The square Poincare norm of each mapped input question word granularity semantic coding vector in the sequence of mapped input question word granularity semantic coding vectors is calculated as the semantic significance description factor to obtain the sequence of input question word granularity semantic significance description factors.

4. The content management method based on big data according to claim 3 is characterized in that: Based on the sequence of the input question word granular semantic significance description factors, hierarchical mask modulation is performed on the sequence of the input question word granular semantic encoding vectors to obtain the sequence of the input question word granular semantic enhancement encoding vectors, including: Performing hierarchical mask modulation on the sequence of input question word granular semantic significance description factors to obtain a sequence of input question word granular semantic significance modulation weights; Taking each input question word granularity semantic significance modulation weight in the sequence of input question word granularity semantic significance modulation weights as a weight, weighted modulation is performed on the corresponding input question word granularity semantic coding vector in the sequence of input question word granularity semantic coding vectors to obtain the sequence of input question word granularity semantic reinforcement coding vectors.

5. The content management method based on big data according to claim 4 is characterized in that: The sequence of input question word granular semantic significance description factors is subjected to hierarchical mask modulation to obtain a sequence of input question word granular semantic significance modulation weights, including: The sequence of input question word granular semantic significance description factors is input and normalized based on a softmax function to obtain a sequence of normalized input question word granular semantic significance description factors; The sequence of normalized input question word granular semantic significance description factors is input into a multi-level masking function to obtain a sequence of input question word granular semantic significance modulation weights.

6. The content management method based on big data according to claim 5 is characterized in that: Inputting the sequence of normalized input question word granular semantic significance description factors into a multi-level masking function to obtain the sequence of input question word granular semantic significance modulation weights, including: Setting a first mask threshold and a second mask threshold, wherein the second mask threshold is twice the first mask threshold; Each normalized input problem word granularity semantic significance description factor in the sequence of normalized input problem word granularity semantic significance description factors is compared with the first mask threshold and the second mask threshold. If the normalized input problem word granularity semantic significance description factor is greater than the second mask threshold, it is weighted and amplified; if the normalized input problem word granularity semantic significance description factor is greater than the first mask threshold and less than or equal to the second mask threshold, it remains unchanged; if the normalized input problem word granularity semantic significance description factor is less than or equal to the first mask threshold, it is weighted and suppressed, thereby obtaining a sequence of input problem word granularity semantic significance modulation weights.

7. The content management method based on big data according to claim 6 is characterized in that: Performing sentiment recognition on the sequence of the input question word granular semantic enhancement encoding vectors to obtain a sentiment type label encoding vector, including: Cascading the sequence of the input question word granular semantic enhancement encoding vectors to obtain the input question global semantic encoding vector; The input question global semantic encoding vector is input into a classifier-based emotion recognition module to obtain the emotion type label encoding vector.

8. The content management method based on big data according to claim 7 is characterized in that: Generating a prompt word based on the sequence of the emotion type label encoding vector and the input question word granularity semantic enhancement encoding vector to obtain an input prompt word for the input question includes: Cascading the input question global semantic encoding vector and the sentiment type label encoding vector to obtain an input question semantic-sentiment cascade encoding vector; The input question semantic-sentiment cascade encoding vector is input into a decoder-based prompt word generator to obtain an input prompt word for the input question.

9. A content management system based on big data, characterized in that: include: A word-granularity semantic encoding module, used for performing word-granularity semantic encoding on the input question to obtain a sequence of word-granularity semantic encoding vectors of the input question; A semantic explicit modeling module, used for performing semantic explicit modeling on the sequence of the input question word granular semantic encoding vectors to obtain a sequence of input question word granular semantic enhanced encoding vectors; An input question emotion recognition module is used to perform emotion recognition on the sequence of the input question word granularity semantic enhancement encoding vectors to obtain an emotion type label encoding vector; The prompt word generation module is used to generate prompt words based on the sequence of the emotion type label encoding vector and the input question word granularity semantic enhancement encoding vector to obtain the input prompt word of the input question.

Citation Information

Patent Citations

  • Content management method and device based on cue word question and answer interaction, equipment and medium

    CN116860935A