A natural language processing scoring method, system, device and medium based on regulatory focus theory

By constructing a corpus of regulatory focus theory and an improved SnowNLP model, the promotion and defense tendencies of entrepreneurs' language are calculated, which solves the customization and personalization needs of existing tools in entrepreneurial communication, realizes more accurate communication analysis and evaluation tools, and enhances the support capabilities of entrepreneurial support agencies.

CN119128153BActive Publication Date: 2025-09-16LANZHOU UNIV OF FINANCE & ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411220431.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-09-16
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

Existing natural language processing tools lack customization and depth when dealing with entrepreneurial communication and business strategy analysis. They have difficulty identifying complex business terms and analyzing the language style of entrepreneurs, and cannot meet the personalized communication needs of entrepreneurs. Entrepreneurship support agencies lack effective evaluation and training tools.

Method used

A corpus based on regulatory focus theory was constructed, the SnowNLP model was improved to the RegFocusNLP model, the Optuna framework was used for parameter optimization, the promotion tendency and defense tendency values ​​were calculated, the regulatory focus index was comprehensively calculated, and a natural language processing scoring system based on regulatory focus theory was provided.

Benefits of technology

It improves the efficiency and effectiveness of entrepreneurial communication, helps entrepreneurs understand and adjust communication strategies, enhances the evaluation and training capabilities of entrepreneurial support agencies, and adapts to different cultures and market needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119128153B_ABST
    Figure CN119128153B_ABST
Patent Text Reader

Abstract

The present invention discloses a natural language processing scoring method, system, device, and medium based on regulatory focus theory, relating to the field of natural language processing technology. The method comprises: constructing a corpus based on regulatory focus theory; the corpus includes entrepreneurial information and corresponding classification labels; constructing an improved SnowNLP model based on the corpus, and optimizing model parameters using the Optuna framework to obtain a RegFocusNLP model; the improved SnowNLP model is constructed based on the analysis function of the promotion tendency and defense tendency of the regulatory focus theory; using the RegFocusNLP model to calculate the target natural language to obtain two classification values; the two classification values ​​include a promotion tendency value and a defense tendency value; and performing a comprehensive calculation on the two classification values ​​to obtain a final regulatory focus index. The present invention can improve the efficiency and effectiveness of entrepreneurial communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a natural language processing scoring method, system, device and medium based on regulatory focus theory. Background Art

[0002] In recent years, with the rapid development of internet and mobile communications technologies, emerging business models such as the sharing economy, e-commerce, and artificial intelligence have rapidly expanded. However, with increasingly fierce market competition, the challenges faced by entrepreneurs are also increasing. In this environment, communication strategies have become a key factor in their success. However, due to the diversity of cultural and economic backgrounds, how to adapt language to meet the expectations of different investors and consumers has become a pressing issue for entrepreneurs.

[0003] Regulatory focus theory provides a powerful tool for understanding and improving entrepreneurs' communication strategies, but it still has the following flaws and shortcomings:

[0004] Limitations of General-Purpose Natural Language Processing Tools: Current natural language processing (NLP) tools on the market are primarily designed for a broad range of language understanding and generation tasks, such as sentiment analysis, text summarization, and machine translation. These tools perform well when processing general text, but often lack the necessary customization and depth when applied to specific domains, such as entrepreneurial communication and business strategy analysis. For example, these tools may not accurately identify and interpret complex business terms or industry-specific expressions in business plans, nor can they conduct in-depth analysis of entrepreneurs' language styles to identify their regulatory focus tendencies.

[0005] The application of psychological motivation theory in NLP is insufficient: Although regulatory focus theory has been widely studied in psychology and management, and its impact on individual behavior and decision-making processes has been demonstrated, relatively few examples of effectively integrating this theory with technology in natural language processing are available. Existing NLP systems rarely consider the impact of users' psychological motivations and goal-seeking strategies on language use, which limits their effectiveness in business strategy analysis and entrepreneur guidance.

[0006] The personalized needs of entrepreneurs' communication strategies cannot be met: In a rapidly changing market environment, especially in a culturally diverse and rapidly developing country like China, entrepreneurs' communication needs are highly personalized and culturally sensitive. Existing NLP tools often lack a deep understanding of regional cultural differences, industry-specific contexts, and individual communication styles. As a result, they may not provide sufficient support for entrepreneurs to adapt their communication strategies to meet the expectations of different cultures and business environments.

[0007] Inadequate evaluation and training tools for startup support organizations: Incubators, accelerators, and other startup support organizations lack tools to accurately analyze and improve entrepreneurs' communication effectiveness when evaluating and training their resident companies. Existing technologies often fail to provide sufficient quantitative data and qualitative feedback to guide entrepreneurs in tailoring their presentations and business documents to their regulatory focus and market needs. Summary of the Invention

[0008] The purpose of the present invention is to provide a natural language processing scoring method, system, device and medium based on regulatory focus theory, which can improve the efficiency and effectiveness of entrepreneurial communication.

[0009] To achieve the above object, the present invention provides the following solutions:

[0010] A natural language processing scoring method based on regulatory focus theory, comprising:

[0011] Constructing a corpus based on regulatory focus theory; the corpus includes entrepreneurial information and corresponding classification labels; wherein the entrepreneurial information includes the entrepreneur's speech content in entrepreneurial programs, entrepreneurial interviews, and entrepreneurial documents;

[0012] An improved SnowNLP model is constructed based on the corpus, and model parameters are optimized using the Optuna framework to obtain the RegFocusNLP model; the improved SnowNLP model is constructed based on the analysis function of the promotion tendency and defense tendency of the regulatory focus theory;

[0013] Utilizing the RegFocusNLP model to calculate the target natural language, a two-category value is obtained; the two-category value includes a promotion tendency value and a defense tendency value;

[0014] The two classification values ​​are comprehensively calculated to obtain the final adjustment focus index.

[0015] Optionally, the corpus is constructed as follows:

[0016] The speech content of entrepreneurs in entrepreneurial programs, entrepreneurial interviews, and entrepreneurial documents was sorted out, and labeled and classified according to the regulatory focus dictionary of regulatory focus theory. The categories of sentences were marked, and the promotion category was marked as 1 and the defense category was marked as 0. The punctuation marks of the collected text content were deleted, and the collected text was segmented using Jieba word segmenter, and the stop words in the text were removed to complete the construction of the corpus.

[0017] Optionally, the improved SnowNLP model is constructed as follows:

[0018] The sentiment analysis function in the traditional SnowNLP model was modified into an analysis function of promotion tendency and defense tendency based on the regulatory focus theory, and the model was trained based on the corpus, Optuna framework, model training formula and bag-of-words model.

[0019] Optionally, the Optuna framework uses a TPE algorithm for parameter optimization.

[0020] Optionally, a comprehensive calculation is performed on the two classification values ​​to obtain a final adjustment focus index, specifically including:

[0021] Based on the promotion tendency value and the defense tendency value, a normalized interpolation method is used to obtain a preliminary value of the regulatory focus index, and the preliminary value is smoothed using the hyperbolic tangent function tanh to obtain a final regulatory focus index.

[0022] The present invention also provides a natural language processing scoring system based on regulatory focus theory, comprising:

[0023] A corpus construction unit, configured to construct a corpus based on regulatory focus theory; the corpus includes entrepreneurial information and corresponding classification labels; wherein the entrepreneurial information includes entrepreneur speeches in entrepreneurial programs, entrepreneurial interviews, and entrepreneurial documents;

[0024] A model building unit, configured to build an improved SnowNLP model based on the corpus and optimize model parameters using the Optuna framework to obtain a RegFocusNLP model; the improved SnowNLP model is built based on the analysis function of the promotion tendency and defense tendency of the regulatory focus theory;

[0025] a text classification unit, configured to calculate the target natural language using the RegFocusNLP model to obtain two classification values; the two classification values ​​include a promotion tendency value and a defense tendency value;

[0026] The scoring calculation unit is used to perform comprehensive calculation on the two classification values ​​to obtain a final adjustment focus index.

[0027] The present invention also provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the natural language processing scoring method based on the regulatory focus theory.

[0028] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the natural language processing scoring method based on the regulatory focus theory as described above.

[0029] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0030] The present invention discloses a natural language processing scoring method, system, device, and medium based on regulatory focus theory. The method includes constructing a corpus based on regulatory focus theory; the corpus includes entrepreneurial information and corresponding classification labels; constructing an improved SnowNLP model based on the corpus, and optimizing model parameters using the Optuna framework to obtain a RegFocusNLP model; the improved SnowNLP model is constructed based on the analysis function of the promotion tendency and defense tendency of the regulatory focus theory; using the RegFocusNLP model to calculate the target natural language to obtain two classification values; the two classification values ​​include a promotion tendency value and a defense tendency value; and performing a comprehensive calculation on the two classification values ​​to obtain a final regulatory focus index. The present invention can improve the efficiency and effectiveness of entrepreneurial communication. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0032] Figure 1 Schematic diagram of the flow of the natural language processing scoring method based on the regulatory focus theory of the present invention;

[0033] Figure 2 Schematic diagram of the scoring process in this embodiment;

[0034] Figure 3 This is a word cloud diagram of the promotion tendency dictionary in this embodiment;

[0035] Figure 4 This is a word cloud diagram of the defense tendency dictionary in this embodiment. DETAILED DESCRIPTION

[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0037] The purpose of the present invention is to provide a natural language processing scoring method, system, device and medium based on regulatory focus theory, which can improve the efficiency and effectiveness of entrepreneurial communication.

[0038] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0039] like Figure 1 As shown, the present invention provides a natural language processing scoring method based on regulatory focus theory, comprising:

[0040] Step 100: Construct a corpus based on regulatory focus theory; the corpus includes entrepreneurial information and corresponding classification labels; wherein the entrepreneurial information includes the entrepreneur's speech content in entrepreneurial programs, entrepreneurial interviews, and entrepreneurial documents;

[0041] Step 200: constructing an improved SnowNLP model based on the corpus, and optimizing model parameters using the Optuna framework to obtain a RegFocusNLP model; the improved SnowNLP model is constructed based on the analysis function of the promotion tendency and defense tendency of the regulatory focus theory;

[0042] Step 300: Calculate the target natural language using the RegFocusNLP model to obtain two-category values; the two-category values ​​include a promotion tendency value and a defense tendency value;

[0043] Step 400: Perform comprehensive calculation on the two classification values ​​to obtain a final adjustment focus index.

[0044] Based on the above technical solution, the following Figure 2-Figure 4 The embodiment shown.

[0045] In this embodiment, based on the promotion and prevention hypotheses of regulatory focus theory, we manually organize the speech content of entrepreneurs in entrepreneurial programs, interviews, and documents, classify it according to the regulatory focus dictionary of regulatory focus theory, and establish a regulatory focus theory corpus. The topics collected for entrepreneurial programs and interviews are the entrepreneurs' responses to the host's questions and the speeches they give when introducing their entrepreneurial projects. These are collected and organized using video-to-text technology. The topics collected for entrepreneurial documents are the entrepreneurs' introductions to their entrepreneurial projects and the speeches of entrepreneurs in meeting minutes in entrepreneurial documents. Optical character recognition (OCR) based on deep learning technology (CRNN+CTC) is used to convert non-digital text into digitized text. Later, the aggregated data is manually cleaned, and the sentence categories are marked, with promotion categories marked as 1 and prevention categories as 0. Punctuation is removed from the collected text content. The collected text is segmented using the Jieba word segmenter, and meaningless stop words such as conjunctions and adverbs are removed. This completes the establishment of a regulatory focus theory corpus.

[0046] SnowNLP is a Python library for processing Chinese texts. Its design was inspired by another text processing library, TextBlob. SnowNLP facilitates users to perform Chinese natural language processing tasks, including sentiment analysis, text summarization, word segmentation, etc. Since SnowNLP is not a natural language processing tool based on the regulatory focus theory, this embodiment replaces its corpus and modifies its functions. Specifically, based on the sentiment analysis function of SnowNLP, its sentiment analysis function is based on the naive Bayes classifier. This embodiment modifies its sentiment analysis function into an analysis function based on the promotion tendency and defense tendency of the regulatory focus theory. The specific model training formula is as follows:

[0047]

[0048] Where P(c|x) is the probability that a given text (x) belongs to category (c) (promotion or defense). P(x|c) is the probability that text (x) appears in category (c), P(c) is the prior probability that any text belongs to category (c), and P(x) is the probability that text (x) appears.

[0049] Feature representation typically uses the bag-of-words model, where the importance of each word is adjusted using TF-IDF. TF-IDF combines the term frequency (TF) and inverse document frequency (IDF) statistics to assess the importance of a word in a text.

[0050]

[0051] TF-DF(t,d,D)=TF(t,d)×IDF(t,D)

[0052] Among them, f t,d is the frequency of word (t) appearing in document (d), and N is the total number of documents in the document set (D).

[0053] Optuna is an automated machine learning (AutoML) framework specifically designed to automate the hyperparameter optimization process of machine learning models. Its working principle is based on efficient search strategies such as Bayesian optimization, genetic algorithms, and TPE (Tree-structured Parzen Estimator). Through these strategies, Optuna is able to explore the optimal combination of hyperparameters to improve model performance. Specifically, Optuna performs hyperparameter optimization by defining an optimization objective (objective function), that is, a metric (such as error rate, accuracy, etc.) that the user wishes to minimize or maximize. The user defines this objective by writing a function, which includes the model training and validation process and returns a score to be optimized.

[0054] Among them, Optuna mainly uses the TPE algorithm for parameter optimization. The algorithm uses the idea of ​​Bayesian optimization to predict the probability distribution of model performance under given hyperparameters by constructing a probability model. That is, the TPE algorithm divides the hyperparameter space into two parts: parameters with better performance and parameters with worse performance. These two parts of parameters are modeled separately to form two probability models (l(x)) and (g(x)), where (l(x)) represents the probability density function of the parameters that make the objective function perform better, and (g(x)) represents the probability density function of the parameters that make the objective function perform worse. When new parameter sampling is performed, TPE samples parameters from (l(x)) instead of uniformly sampling from the original or entire parameter space. The sampling strategy is based on the following optimization formula:

[0055]

[0056] Here x * is the next parameter point determined during the evaluation process. Specifically, the point that maximizes the ratio of (l(x)) to (g(x)) is selected as the next sampling point. After each experiment, (l(x)) and (g(x)) are updated based on the experimental results. This update process relies on the premise of updating the posterior probability distribution in Bayesian theory, which continuously adjusts the parameters of the probability model based on new experimental results. Through this iterative process, TPE can gradually narrow the search scope, focusing on those hyperparameters most likely to achieve better performance, thereby achieving more efficient hyperparameter optimization.

[0057] By replacing the modified SnowNLP model and using Optuna to automatically optimize its parameters, a natural language processing model based on the regulatory focus theory suitable for the entrepreneurial field is constructed, namely the RegFocusNLP model. Based on this model, the promotion tendency value and the defense tendency value are calculated, and the classification of natural language text is completed. The promotion tendency value and the defense tendency value are the possibilities of the RegFocusNLP model classifying it as a promotion tendency or a defense tendency. Based on the obtained promotion tendency value and defense tendency value, the normalized interpolation method is used to obtain the preliminary value of the regulatory focus index, and the obtained preliminary value is smoothed by the Sigmoid adjustment method, that is, the hyperbolic tangent function tanh is used, which will limit the result to between -1 and 1 and provide a smoother output when the input value is small (that is, the promotion and defense scores are similar). The formula is expressed as:

[0058]

[0059] In general, by replacing the SnowNLP corpus and modifying the classification tasks and classification targets, and further optimizing the naive Bayes classifier used by SnowNLP through Optuna parameter optimization, a more accurate natural language classification task is achieved, and a natural language processing model (RegFocusNLP) suitable for the regulatory focus theory is constructed. The promotion tendency value and the defense tendency value obtained are comprehensively calculated, and the regulatory focus index is calculated based on the normalization method. The Sigmoid adjustment method is used to smooth the calculated value to obtain the final regulatory focus index, and finally a natural language processing scoring system RegFocusNLP System based on the regulatory focus theory is constructed. This patent not only improves the accuracy of natural language classification tasks, but also provides entrepreneurs with a new tool to help them better understand and adjust their business strategies and communication methods. In addition, this system also provides an effective evaluation and training tool for entrepreneurial support agencies, further enhancing their ability to support entrepreneurs.

[0060] Therefore, the present invention has the following beneficial effects:

[0061] (1) By combining regulatory focus theory with natural language processing technology, this system specifically classifies and scores promotional and defensive tendencies. Compared to traditional sentiment analysis tools (which typically only distinguish between positive and negative sentiment), this system can provide more detailed and targeted analysis. This precise classification can help entrepreneurs better understand the impression their texts (such as business plans and investment presentations) may leave on investors or potential partners, allowing them to make necessary adjustments to align their business goals with market demand.

[0062] (2) By accurately identifying and analyzing regulatory focus trends in entrepreneurship-related texts, the present invention provides a powerful decision-making support tool. Entrepreneurs can use this tool to optimize their communication strategies and business plans, ensuring that they are consistent with the company's long-term goals and market positioning. In addition, entrepreneurship support organizations such as incubators and accelerators can also use this system to evaluate and improve the performance of their resident companies, thereby allocating resources and support more effectively.

[0063] (3) Considering that the expression of promotion and defense focus may vary across different cultural contexts, the application of this system can help entrepreneurs make effective language and strategy adjustments when preparing for international market promotion and cross-cultural communication. Such adjustments can not only reduce cultural misunderstandings but also enhance the attractiveness and competitiveness of entrepreneurial projects in the global market. By sensitively addressing and reflecting the needs of different markets and cultures, entrepreneurs can more successfully expand internationally.

[0064] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0065] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A natural language processing scoring method based on regulatory focus theory, characterized in that: include: Constructing a corpus based on regulatory focus theory; the corpus includes entrepreneurial information and corresponding classification labels; wherein the entrepreneurial information includes the entrepreneur's speech content in entrepreneurial programs, entrepreneurial interviews, and entrepreneurial documents; An improved SnowNLP model is constructed based on the corpus, and model parameters are optimized using the Optuna framework to obtain the RegFocusNLP model; the improved SnowNLP model is constructed based on the analysis function of the promotion tendency and defense tendency of the regulatory focus theory; Utilizing the RegFocusNLP model to calculate the target natural language, a two-category value is obtained; the two-category value includes a promotion tendency value and a defense tendency value; Performing a comprehensive calculation on the two classification values ​​to obtain a final adjustment focus index; The construction process of the corpus is as follows: The speech content of entrepreneurs in entrepreneurial programs, entrepreneurial interviews, and entrepreneurial documents was sorted and categorized according to the regulatory focus dictionary of regulatory focus theory. The categories of sentences were marked, and the promotion category was marked as 1 and the defense category was marked as 0. The punctuation marks in the collected text content were deleted, and the collected text was segmented using the Jieba word segmenter. The stop words in the text were removed to complete the construction of the corpus. The construction method of the improved SnowNLP model is: The sentiment analysis function in the traditional SnowNLP model was modified to an analysis function based on the promotion tendency and defense tendency of regulatory focus theory. The model was trained based on the corpus, Optuna framework, model training formula and bag-of-words model. The two classification values ​​are comprehensively calculated to obtain the final adjustment focus index, which specifically includes: Based on the promotion tendency value and the defense tendency value, a normalized interpolation method is used to obtain a preliminary value of the regulatory focus index, and the preliminary value is smoothed using the hyperbolic tangent function tanh to obtain a final regulatory focus index.

2. The natural language processing scoring method based on regulatory focus theory according to claim 1, characterized in that: The Optuna framework uses the TPE algorithm for parameter optimization.

3. A natural language processing scoring system based on regulatory focus theory, applying the method according to any one of claims 1-2, characterized in that: include: A corpus construction unit, configured to construct a corpus based on regulatory focus theory; the corpus includes entrepreneurial information and corresponding classification labels; wherein the entrepreneurial information includes entrepreneur speeches in entrepreneurial programs, entrepreneurial interviews, and entrepreneurial documents; A model building unit, configured to build an improved SnowNLP model based on the corpus and optimize model parameters using the Optuna framework to obtain a RegFocusNLP model; the improved SnowNLP model is built based on the analysis function of the promotion tendency and defense tendency of the regulatory focus theory; a text classification unit, configured to calculate the target natural language using the RegFocusNLP model to obtain two classification values; the two classification values ​​include a promotion tendency value and a defense tendency value; The scoring calculation unit is used to perform comprehensive calculation on the two classification values ​​to obtain a final adjustment focus index.

4. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the natural language processing scoring method based on the regulatory focus theory according to any one of claims 1-2.

5. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed by a processor, implements the natural language processing scoring method based on the regulatory focus theory as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Optimizing dialogue policy decisions for digital assistants using implicit feedback

    CN110637339A

  • Non-perpetual propagation group sentiment analysis method based on deep learning and analysis system thereof

    CN117540740A