Big data-based copywriting semantic recognition system and method

By designing a copy-text semantic recognition system based on big data, the problems of data quality problems and the impact of false data in the existing system are solved, and more efficient and accurate copy-text semantic recognition is achieved.

CN120218083AInactive Publication Date: 2025-06-27GUANGDONG BOJIN INFORMATION TECHNOLOGY GROUP CO LTD

Patent Information

Application Number
CN202510682602.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing copywriting semantic recognition systems have the problem of uneven data quality, and there are often duplicate and malformed data in the data, and the lack of an effective data verification mechanism, resulting in inaccuracy and inefficiency of identification.

Method used

A copywriting semantic recognition system based on big data is designed, including data acquisition module, preliminary cleaning, semantic recognition model training and execution module. The system automatically collects copywriting data, performs preliminary cleaning and removes duplicate and format error data, and uses deep learning algorithms to train semantic recognition models. At the same time, a data verification unit is introduced to ensure the authenticity of the data through multi-mode verification.

Benefits of technology

Through the support of big data storage modules and the application of deep learning algorithms, the accuracy and efficiency of semantic recognition are improved. The multi-mode verification mechanism of the data verification unit effectively reduces the impact of false data and ensures the reliability of the identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218083A_ABST
    Figure CN120218083A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, and discloses a copywriting semantic recognition system and method based on big data, and the system comprises a data collection module which collects copywriting data from a plurality of data sources, automatically collects new data at a set time interval, and carries out the preprocessing of the collected data, repeating data and format error data are removed; the data source comprises a webpage, a social media platform and a document database; the big data storage module stores the preliminarily cleaned copywriting data; the semantic recognition model training module is used for training a semantic recognition model by using the data stored in the big data storage module; the semantic recognition execution module is connected with the semantic recognition model training module and the big data storage module, a trained semantic recognition model is used for conducting semantic recognition on newly-input copywriting, a recognition result is output, and the recognition result at least comprises subject classification, emotional tendency judgment and key semantic information extraction of the copywriting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a copywriting semantic recognition system and method based on big data. Background Art

[0002] In the current era of digital information explosion, copywriting data shows an explosive growth trend. Copywriting widely exists in web pages, social media platforms, and various document databases. These copywritings contain rich information, such as product promotion copywriting, news reports, user comments, etc. Conducting semantic recognition on these copywritings to obtain their topic classification, sentiment tendency, and key semantic information has important commercial and social value.

[0003] However, there are still some defects in the existing copywriting semantic recognition systems. For example: Currently, although there are some methods to collect copywriting data from different data sources, there are problems with uneven data quality; on the one hand, due to numerous data sources and the lack of a unified specification, the collected data often contains a large amount of duplicate data and data with format errors. These useless data will not only occupy a large amount of storage resources but also interfere with subsequent semantic recognition work, reducing the accuracy and efficiency of recognition. On the other hand, the existing data collection process often lacks an effective data verification mechanism, making it difficult to ensure the authenticity of the collected data. False data may mislead the semantic recognition results, causing deviations in the decisions made based on these results. Summary of the Invention

[0004] The purpose of the present invention is to provide a copywriting semantic recognition system and method based on big data to solve the above technical problems.

[0005] The purpose of the present invention can be achieved through the following technical solutions: A copywriting semantic recognition system based on big data, comprising: A data collection module, used to collect copywriting data from multiple data sources, automatically collect new data at a set time interval, and preprocess the collected copywriting data, including preliminary cleaning to remove duplicate data and data with format errors; the data sources include web pages, social media platforms, and document databases; A big data storage module, used to store the copywriting data after preliminary cleaning; A semantic recognition model training module, which uses the copywriting data stored in the big data storage module to train a semantic recognition model; A semantic recognition execution module, connected to the semantic recognition model training module and the big data storage module, uses the trained semantic recognition model to perform semantic recognition on newly input copywriting, and outputs recognition results. The recognition results at least include the topic classification of the copywriting, sentiment tendency judgment, and key semantic information extraction.

[0006] As a further technical solution, the semantic recognition model training module adopts deep learning algorithms, including the recurrent neural network (RNN) and its variant, the long short-term memory network (LSTM), and the model based on the Transformer architecture.

[0007] As a further technical solution, the data acquisition module further includes a data verification unit, which is used to verify the authenticity of the collected copywriting data, and eliminate or retain the collected copywriting data according to the verification result of the authenticity; the authenticity verification includes a first verification mode and a second verification mode.

[0008] As a further technical solution, the working process of the first verification mode is as follows: Verify the collected copywriting data according to a preset rule set, where the rule set includes a format rule, a content rationality rule set, and a source reliability rule; Process to obtain the format score of the current copywriting data 、content rationality score and source reliability score ; Calculate the first authenticity score through the formula: ; Where , is the number of illegal characters in the copywriting, is the total number of characters in the copywriting; , is the number of unreasonable key information in the copywriting, is the total number of content key information in the copywriting; , is the historical error rate of the source of the current copywriting data, is the source reliability standard score.

[0009] As a further technical solution, the working process of the second verification mode is as follows: Use a pre-trained authenticity verification model to verify the collected copywriting data; Where the expression of the authenticity verification model is: ; The authenticity verification model takes the feature vector of the collected copywriting data as input, where the feature vector includes the text feature, source feature, and context feature of the copywriting; the output is the second authenticity score that the current copywriting data is real data, is the bias term, is the weight coefficient of the feature ,​ is the exponential coefficient.

[0010] As a further technical solution, the process of rejecting or retaining the collected copywriting data according to the verification result of authenticity is as follows: The first authenticity score calculated , the second authenticity score are substituted into the following formula: ; The authenticity comprehensive index is calculated; , are weight factors, determined based on historical data analysis, , are respectively the maximum and minimum values of the historical second authenticity score; When , it is initially determined that the current copywriting data is authentic data, and then the first authenticity score calculated is compared with the preset first threshold interval , and the second authenticity score calculated is compared with the preset second authenticity threshold for further verification; When , it is determined that the current copywriting data is unauthentic data and is rejected.

[0011] As a further technical solution, the process of comparing the first authenticity score calculated with the preset first threshold interval is as follows: If > , it is determined that the current copywriting data passes the verification of the rule set, is retained as authentic data, and is verified through the second verification mode; If < , it is determined that the current copywriting data is unauthentic data and is rejected; If , the current copywriting data is marked and verified through the second verification mode.

[0012] As a further technical solution, the process of comparing the second authenticity score calculated with the preset second authenticity threshold is as follows: If , it is determined that the current copywriting data is authentic data and is retained; If , then it is determined that the current copywriting data is untrue data and is excluded.

[0013] A method for identifying the semantics of copywriting based on big data, the method comprising the following steps: S1. Collect copywriting data from multiple data sources, automatically collect new data at a set time interval, and preprocess the collected data, including preliminary cleaning to remove duplicate data and data with format errors; the data sources include web pages, social media platforms, and document databases; S2. Store the copywriting data after preliminary cleaning; S3. Use the data stored in the big data storage module to train a semantic recognition model; S4. Connect to the semantic recognition model training module and the big data storage module, use the trained semantic recognition model to perform semantic recognition on newly input copywriting, and output the recognition result, where the recognition result at least includes the theme classification of the copywriting, the judgment of the emotional tendency, and the extraction of key semantic information.

[0014] Advantages of the present invention: In the present invention, the big data storage module ensures the safe storage and convenient call of data, provides stable data support for model training and recognition. For example, an e-commerce platform stores a large amount of product copywriting data, which is convenient for analysis at any time; the semantic recognition model training module uses deep learning algorithms to improve the model's recognition ability, making the theme classification, emotional tendency judgment, and extraction of key semantic information more accurate. Taking the analysis of social media comments as an example, it can accurately grasp the emotions and key demands of users; the data verification unit verifies the authenticity of data in multiple modes, reduces the impact of false data, and provides a reliable basis for subsequent decisions. In the analysis of copywriting in the financial field, it can avoid investment decision-making mistakes caused by false data. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The present invention will be further described below with reference to the accompanying drawings.

[0016] Figure 1 It is a structural block diagram of the copywriting semantic recognition system based on big data of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] Please refer to Figure 1 As shown, the present invention is a copywriting semantic recognition system based on big data, including: A data collection module for collecting copywriting data from multiple data sources, automatically collecting new data at set time intervals, and preprocessing the collected copywriting data, including preliminary cleaning to remove duplicate data and data with formatting errors; the data sources include web pages, social media platforms, and document databases; A big data storage module for storing the copywriting data after preliminary cleaning; A semantic recognition model training module that uses the copywriting data stored in the big data storage module to train a semantic recognition model; the training process is a prior art and can achieve semantic recognition, so it will not be elaborated here; A semantic recognition execution module, connected to the semantic recognition model training module and the big data storage module, uses the trained semantic recognition model to perform semantic recognition on newly input copywriting, and outputs the recognition results, where the recognition results at least include the theme classification of the copywriting, sentiment tendency judgment, and extraction of key semantic information.

[0019] In this embodiment, the big data storage module ensures the secure storage and convenient invocation of data, providing stable data support for model training and recognition. For example, an e-commerce platform stores a large amount of product copywriting data for easy analysis at any time; the semantic recognition model training module uses deep learning algorithms to improve the model's recognition ability, making the theme classification, sentiment tendency judgment, and extraction of key semantic information more accurate. Taking the analysis of social media comments as an example, it can accurately grasp users' emotions and key demands; the data verification unit verifies the authenticity of data in multiple modes, reducing the impact of false data and providing a reliable basis for subsequent decisions. In the analysis of copywriting in the financial field, it can avoid investment decision-making mistakes caused by false data; Taking an e-commerce platform as an example, the data collection module regularly collects copywriting data from the product details page: web page data source, user evaluation area: social media platform data source, and internal document library: document database data source. It automatically collects data at 2 am every day, and removes duplicate evaluations and data with garbled formats after collection; the big data storage module stores the cleaned data by product category; the semantic recognition model training module uses this data to train the model, and the semantic recognition execution module recognizes the copywriting of newly listed products, judging the theme classification such as electronic products, clothing, the sentiment tendency such as users' satisfaction with the product, and the key semantic information such as product features and usage experiences.

[0020] The semantic recognition model training module uses deep learning algorithms, including the recurrent neural network (RNN) and its variant, the long short-term memory network (LSTM), and models based on the Transformer architecture. For example, in the scenario of public opinion monitoring, the LSTM network is used to train the semantic recognition model. After the collected copywriting data such as news and social media posts are tokenized and vectorized, they are input into the model. The model learns the long-term and short-term dependencies in the text. After training, it can identify new public opinion copywriting, such as judging whether the theme of a certain enterprise-related public opinion is product quality problems or corporate image promotion, and the sentiment tendency is positive, negative or neutral, and extract key semantic information such as problem focus and public opinion orientation. By using deep learning algorithms, the model can better capture the semantic features of the text, improve the semantic recognition accuracy, help enterprises keep abreast of public opinion dynamics in a timely manner, and formulate response strategies.

[0021] The data acquisition module further includes a data verification unit for verifying the authenticity of the collected copywriting data and deleting or retaining the collected copywriting data according to the verification result of the authenticity. The authenticity verification includes a first verification mode and a second verification mode.

[0022] In this embodiment, a data verification unit is introduced. The data verification unit uses two verification modes to verify the authenticity of the collected copywriting data. The first verification mode evaluates the data from three dimensions: format, content rationality, and source reliability according to a preset rule set. By calculating the format score, content rationality score, and source reliability score, the first authenticity score is obtained comprehensively. The second verification mode uses a pre-trained authenticity verification model, takes the text, source, and context feature vectors of the copywriting data as inputs, and outputs the second authenticity score. Finally, the two scores are substituted into the comprehensive index formula, and it is judged whether the data is true according to the calculation result combined with a preset threshold, and the data is decided to be deleted or retained. The data verification unit ensures the authenticity of the collected data through two verification modes, which is crucial for the accuracy and reliability of subsequent semantic recognition.

[0023] The working process of the first verification mode is as follows: Verify the collected copywriting data according to the preset rule set, and the rule set includes format rules, content rationality rule sets, and source reliability rules; Process to obtain the format score of the current copywriting data 、content rationality score and source reliability score ; Through the formula: Calculate to obtain the first authenticity score ; Through reflects the combined influence of format and content rationality on data authenticity, and through the exponential function in the form of adjusting the comprehensive score according to the reliable source score; when the reliable source score is high, the denominator approaches 1, and the first authenticity score depends more on the format and content rationality scores; when is low, the denominator becomes larger, suppressing the overall score; thus, the evaluation results of the three dimensions can be integrated into a comprehensive score, comprehensively reflecting the authenticity degree of the data verified based on the rule set, providing a quantitative basis for subsequent data screening; Among them, , is the number of illegal characters in the copywriting, and the acquisition method is as follows: according to the pre-set legal character set, each character in the copywriting is checked; the legal character set can be defined according to specific requirements. For example, in Chinese text processing, the legal character set can include Chinese characters, common punctuation marks, numbers, and English letters, etc.; traverse the characters in the copywriting, if a certain character is not in the legal character set, then add 1 to the number of illegal characters, is the total number of characters in the copywriting; this formula calculates the proportion of illegal characters in the total characters, and uses the complement of the proportion as the format score. The higher the proportion, the lower the format score, which intuitively reflects the standardization of the copywriting format, so as to quickly identify data with serious format problems, such as copywriting with excessive garbled characters. Such data may interfere with subsequent processing. Through the format score, low-quality data can be initially screened out to ensure that the data format meets the requirements of subsequent processing, improving the efficiency and accuracy of data processing.

[0024] , is the number of unreasonable key information in the copywriting, and the acquisition method is as follows: first, key information needs to be defined. Key information can be the core concepts, events, data, etc. involved in the copywriting, and then judge whether the key information is reasonable according to the pre-set rules or knowledge base. For example, for a news report, the key information includes the time, place, people, and event content when the event occurred. Judge the reasonableness of the key information by checking whether the time is logical, whether the place exists, and whether the person information is accurate, etc.; if a certain key information does not conform to the content in the rule or knowledge base, then add 1 to the number of unreasonable key information, is the total number of content key information in the copywriting, and the acquisition method is as follows: through natural language processing technologies, such as named entity recognition, information extraction, etc. to identify the key information in the copywriting, and then calculate its quantity; The above formula measures the content rationality by the proportion of unreasonable key information in the total key information. The lower the proportion, the higher the score. It can evaluate data from the content logic level, such as judging whether the event description in the news copy conforms to common sense, whether the time and place are accurate, etc., so as to exclude data with logical errors or unreasonable expressions in the content, ensure the reliability of the data entering the subsequent process, and avoid affecting the semantic recognition result due to incorrect information; , is the historical error rate of the current copy data source, and the acquisition method is: record the error situations in the data provided by each data source in the past; for each data source, count the total number of data it provides and the number of data with errors among them, then the historical error rate , is the source reliability standard score; this formula adjusts the source reliability score according to the historical error situation of the data source. The higher the historical error rate, the lower the source reliability score. For example, when collecting data, if a certain website often publishes incorrect information, the source reliability score of the corresponding data is low, so as to consider the credibility of the data source, preferentially retain data from reliable sources, reduce the risk of data errors caused by unreliable sources, and improve the overall quality of the data.

[0025] The working process of the second verification mode is as follows: Use a pre-trained authenticity verification model to verify the collected copy data; Among them, the expression of the authenticity verification model is: ; The authenticity verification model takes the feature vector of the collected copy data as input, where the feature vector includes the text feature, source feature and context feature of the copy; the output is the second authenticity score indicating that the current copy data is real data , is the bias term, is the weight coefficient of the feature , determined based on historical data analysis, is the exponential coefficient.

[0026] In this embodiment, through the feature vector as input, is the bias term, is the weight coefficient of the feature , ; Based on the principle of logistic regression, by performing weighted summation on the input feature vector and passing through Function transformation outputs the second authenticity score between 0 and 1. Using a machine learning model, it comprehensively judges the data authenticity from multiple feature dimensions, which is more comprehensive and intelligent than single-rule judgment, can conduct a more in-depth and comprehensive evaluation of data authenticity, mine potential authenticity features of data, and improve the accuracy of data verification.

[0027] The process of rejecting or retaining the collected copywriting data according to the verification result of authenticity is as follows: Substitute the calculated first authenticity score and the second authenticity score into the following formula: ; Calculate the comprehensive authenticity index ; , are weight factors determined based on historical data analysis, , are respectively the maximum and minimum values of the historical second authenticity scores; When , initially judge that the current copywriting data is authentic data, and then compare the calculated first authenticity score with the preset first threshold interval , and compare the calculated second authenticity score with the preset second authenticity threshold for further verification; When , judge that the current copywriting data is unauthentic data and reject it.

[0028] In this embodiment, the first authenticity score and the second authenticity score are standardized and then weighted and summed to obtain the comprehensive authenticity index. The standardization process makes the scores under different scoring systems comparable, and the weight factors determine the relative importance of the two scores in the comprehensive judgment according to historical data, so as to be able to integrate the results of the two verification modes, give a more comprehensive and objective authenticity judgment index, reduce the misjudgment that may occur in a single verification mode, and improve the reliability of data verification.

[0029] The process of comparing the calculated first authenticity score with the preset first threshold interval is as follows: If > , then judge that the current copywriting data passes the verification of the rule set, is retained as authentic data, and is verified through the second verification mode; If < , then it is determined that the current copywriting data is untrue data and is excluded; If , then mark the current copywriting data and verify it through the second verification mode.

[0030] The calculated second authenticity score is compared with the preset second authenticity threshold The process is as follows: If , then it is determined that the current copywriting data is authentic data and is retained; If , then it is determined that the current copywriting data is untrue data and is excluded.

[0031] A copywriting semantic recognition method based on big data, the method includes the following steps: S1. Collect copywriting data from multiple data sources, automatically collect new data at a set time interval, and preprocess the collected data, including preliminary cleaning to remove duplicate data and data with format errors; the data sources include web pages, social media platforms, and document databases; S2. Store the copywriting data after preliminary cleaning; S3. Use the data stored in the big data storage module to train a semantic recognition model; S4. Connect to the semantic recognition model training module and the big data storage module, use the trained semantic recognition model to perform semantic recognition on newly input copywriting, and output the recognition result, the recognition result at least includes the theme classification of the copywriting, sentiment tendency judgment, and key semantic information extraction.

[0032] It should be noted that: the calculation formulas and each parameter participating in the operation in the present invention have been pre-dimensionless processed, and the process of dimensionless processing is well-known in the industry and will not be described here.

[0033] The above has described an embodiment of the present invention in detail, but the above content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application should still fall within the scope covered by the patent of the present invention.

Claims

1. A copywriting semantic recognition system based on big data, characterized in that, Including: A data acquisition module, which is used to collect copywriting data from multiple data sources, automatically collect new data at a set time interval, and preprocess the collected copywriting data, including preliminary cleaning to remove duplicate data and data with format errors; The data sources include web pages, social media platforms, and document databases; A big data storage module, which is used to store the copywriting data after preliminary cleaning; A semantic recognition model training module, which uses the copywriting data stored in the big data storage module to train a semantic recognition model; A semantic recognition execution module, which is connected to the semantic recognition model training module and the big data storage module, uses the trained semantic recognition model to perform semantic recognition on newly input copywriting, and outputs the recognition result, and the recognition result at least includes the theme classification of the copywriting, sentiment tendency judgment, and key semantic information extraction.

2. The copywriting semantic recognition system based on big data according to claim 1, wherein The semantic recognition model training module adopts deep learning algorithms, including the recurrent neural network RNN and its variant long short-term memory network LSTM, and models based on the Transformer architecture.

3. The copywriting semantic recognition system based on big data according to claim 2, wherein The data acquisition module further includes a data verification unit, which is used to verify the authenticity of the collected copywriting data, and eliminate or retain the collected copywriting data according to the verification result of the authenticity; the authenticity verification includes a first verification mode and a second verification mode.

4. The copywriting semantic recognition system based on big data according to claim 3, wherein, The working process of the first verification mode is as follows: Verify the collected copywriting data according to a preset rule set, and the rule set includes format rules, content rationality rules, and source reliability rules; Obtain the format score of the current copywriting data , content rationality score and source reliability score ; Through the formula: The first authenticity score is calculated ; Among them, , is the number of illegal characters in the copywriting, is the total number of characters in the copywriting; , is the number of unreasonable key information in the copywriting, is the total number of content key information in the copywriting; , is the historical error rate of the data source of the current copywriting, is the reliable source standard score.

5. The copywriting semantic recognition system based on big data according to claim 4, wherein The working process of the second verification mode is as follows: Use a pre-trained authenticity verification model to verify the collected copywriting data; Among them, the expression of the authenticity verification model is: ; The authenticity verification model takes the feature vectors of the copywriting data collected as input, where the feature vectors include the text features, source features, and context features of the copywriting; the output is the second authenticity score indicating that the current copywriting data is authentic data , is the bias term, is the weight coefficient of the feature , is the exponential coefficient.

6. The copywriting semantic recognition system based on big data according to claim 1 or 5, characterized in that The process of eliminating or retaining the collected copywriting data according to the verification result of the authenticity is as follows: Substitute the calculated first authenticity score , the second authenticity score into the following formula: ; Calculate the comprehensive authenticity index ; , are weight factors determined based on historical data analysis, , are the maximum and minimum values of the second historical authenticity score respectively; When it is initially determined that the current copywriting data is real data, and then the first authenticity score calculated is compared with the preset first threshold range and the second authenticity score calculated is compared with the preset second authenticity threshold for further verification; When it is determined that the current copywriting data is untrue data, and it is excluded.

7. The copywriting semantic recognition system based on big data according to claim 6, characterized in that The calculated first authenticity score and the preset first threshold interval The process of comparison is as follows: If > , it is determined that the current copywriting data passes the verification of the rule set, retained as real data, and verified through the second verification mode; If < , then it is determined that the current copywriting data is untrue data and is excluded; If , mark the current copywriting data and verify it through the second verification mode.

8. The copywriting semantic recognition system based on big data according to claim 6, characterized in that The calculated second authenticity score is compared with a preset second authenticity threshold in the following process: If , it is determined that the current copywriting data is real data and retained; If , then it is determined that the current copywriting data is untrue data and is excluded.

9. A copywriting semantic recognition method based on big data, which is implemented based on the big data-based copywriting semantic recognition system described in claim 1, and is characterized in that, Including the following steps: S1. Collect copywriting data from multiple data sources, automatically collect new data at a set time interval, and preprocess the collected data, including preliminary cleaning to remove duplicate data and data with format errors; The data sources include web pages, social media platforms, and document databases; S2. Store the copywriting data after preliminary cleaning; S3. Use the data stored in the big data storage module to train a semantic recognition model; S4. Connect to the semantic recognition model training module and the big data storage module, use the trained semantic recognition model to perform semantic recognition on newly input copywriting, and output the recognition result, and the recognition result at least includes the theme classification of the copywriting, sentiment tendency judgment, and key semantic information extraction.

Citation Information

Patent Citations

  • Adaptive network queue data processing device based on deep learning

    CN118551348A

  • Authenticity verification method and device

    WO2024174714A1

Cited By

  • High-performance semantic recognition method

    CN120764553A

  • A high-performance semantic recognition method

    CN120764553B