Book limit word monitoring management method, system and device and storage medium

By updating the limit word database at every preset time in the book limit word monitoring management system, the limit word database is updated and the model is trained, and combined with regular expression search and semantic feature analysis, the problem of inaccurate limit word monitoring in the existing technology is solved, efficient and accurate monitoring management is achieved, and the system efficiency and user experience are improved.

CN120196737AInactive Publication Date: 2025-06-24GUOMAI CULTURE MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510084230.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing limit word monitoring and management methods rely on simple keyword matching and cannot accurately understand text semantics, resulting in misjudgment in complex text content, increasing the workload of manual review and reducing the efficiency and reliability of the monitoring and management system.

Method used

By obtaining the limit words of books from the rule databases of multiple e-commerce platforms at every preset time, updating the local database and training the preset model. Use regular expressions to search the full text, locate the target paragraphs containing the limit words, and use the preset model to perform semantic feature analysis to determine whether the target limit words need to be corrected.

Benefits of technology

It improves the accuracy and flexibility of monitoring, reduces the possibility of false alarms and missed reports, improves processing efficiency, ensures that there is no illegal content in the book details page, reduces legal risks and economic losses, and improves users' reading experience and trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196737A_ABST
    Figure CN120196737A_ABST
Patent Text Reader

Abstract

The invention discloses a book limit word monitoring management method, system and device and a storage medium, and relates to the field of limit word monitoring. In the method, limit words of books are obtained from rule libraries of a plurality of e-commerce platforms every preset time, a local database is updated according to the limit words, and a preset model is trained according to the local database; constructing a regular expression according to the limit word, and performing full-text search on a book detail page according to the regular expression to locate a target paragraph containing the limit word; performing semantic feature analysis on the target paragraph through the preset model to determine whether a target limit word in the target paragraph needs to be corrected or not; and when a target limit word in the target paragraph needs to be corrected, generating a source file index link according to a file path of the target limit word, and sending the source file index link to a preset target. By implementing the technical scheme provided by the invention, the limit words in the book introduction can be effectively monitored and managed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of extreme word monitoring, and specifically relates to a method, system, device and storage medium for monitoring and managing extreme words in books. Background Art

[0002] In recent years, with the rapid development of e-commerce, the problem of the standardization of product descriptions has become increasingly prominent. Especially for the descriptions of cultural products such as books, it is often necessary to strictly comply with relevant regulations such as the Advertising Law and avoid using illegal words such as extreme words.

[0003] The existing extreme word monitoring and management methods mainly rely on simple keyword matching and lack the ability to understand the semantics of text. This method is prone to misjudgment when dealing with complex text content. Especially in the case of rich context information, simply relying on keyword matching cannot accurately judge the actual meaning and compliance of extreme words. This not only increases the workload of manual review but also reduces the efficiency and reliability of the entire monitoring and management system.

[0004] Therefore, how to efficiently and accurately monitor and manage extreme words in book descriptions has become an urgent problem for major e-commerce platforms to solve. Summary of the Invention

[0005] This application provides a method, system, device and storage medium for monitoring and managing extreme words in books, which can deeply understand the text content, avoid mechanical judgment of extreme words, and thus improve the accuracy and flexibility of monitoring.

[0006] 1. In the first aspect of this application, a method for monitoring and managing extreme words in books is provided, which is applied to a monitoring and management platform. The method includes: Obtain extreme words of books from the rule libraries of multiple e-commerce platforms at preset intervals, update the local database according to the extreme words, and train a preset model according to the local database; Construct a regular expression according to the extreme words, and perform a full-text search on the book detail page according to the regular expression to locate the target paragraph containing the extreme words; Perform semantic feature analysis on the target paragraph through the preset model to determine whether the target extreme words in the target paragraph need to be corrected; When the target extreme words in the target paragraph need to be corrected, generate a source file index link according to the file path of the target extreme words, and send the source file index link to a preset target.

[0007] 2. Optionally, the training of the preset model according to the local database includes: Extract a data set from the local database, where the data set includes extreme words and the context information of the extreme words; Clean the dataset, convert the data in the cleaned dataset into numerical feature vectors, and divide the cleaned dataset into a training set and a test set according to a preset ratio; Use a neural network architecture to train the preset model through the training set to learn the feature representation and context dependencies of extreme words until a preset number of iterations is reached or the loss function converges; Use the test set to evaluate the preset model, and adjust the hyperparameters or model structure according to the evaluation results until the model performance meets the requirements.

[0008] 3 Optionally, the full-text search of the book details page according to the regular expression to locate the target paragraph containing the extreme word includes: Traverse the text content in the book details page based on the regular expression to match the candidate paragraphs containing the extreme word; Convert the candidate paragraphs into hash values, and remove duplicates from the candidate paragraphs according to the hash values to select the target paragraphs from the candidate paragraphs.

[0009] 4 Optionally, the semantic feature analysis of the target paragraph through the preset model to determine whether the target extreme word in the target paragraph needs to be corrected includes: Convert the text in the target paragraph into a semantic vector through the preset model, and determine the target feature vector of the target extreme word from the semantic vector; Match the usage of the target extreme word in the context with the standard usage in the preset rule library. When the usage is inconsistent with the standard usage, determine that the target extreme word needs to be corrected; When the usage is consistent with the standard usage, match the target feature vector with the standard semantic vector in the preset rule library. When the target feature vector is inconsistent with the standard semantic vector, determine that the target extreme word needs to be corrected.

[0010] 5 Optionally, the method further includes: Receive the correction result of the preset target for the target extreme word, where the correction result includes rejecting the correction of the first target extreme word and the content after correcting the second target extreme word; Perform negative training on the preset model according to the first target extreme word and the context information of the first target extreme word to reduce the recognition of the first target extreme word by the preset model; Perform positive training on the preset model according to the second target extreme word and the content after correcting the second target extreme word to increase the recognition of the second target extreme word by the preset model.

[0011] 6 Optionally, the negative training of the preset model according to the first target extreme word and the context information of the first target extreme word to reduce the recognition of the first target extreme word by the preset model includes: Adding the first target extreme word and the context information of the first target extreme word as negative samples to the training set; Using the updated training set to retrain the preset model to reduce the recognition rate of the negative samples by the preset model.

[0012] 7 Optionally, the generating of the source file index link according to the file path of the target extreme word includes: Parsing the file path where the target extreme word is located, and extracting key information, where the key information includes the file name, file type, and storage location; Constructing a source file index database according to the key information, where the source file index database includes the unique identifier, path information, and modification time of the source file; Generating a corresponding source file index link according to the unique identifier.

[0013] 8 In a second aspect of the present application, a monitoring and management system for book extreme words is provided, including a collection module, a positioning module, an analysis module, and an execution module, where: The collection module is configured to obtain the extreme words of books from the rule libraries of multiple e-commerce platforms at preset intervals, update the local database according to the extreme words, and train a preset model according to the local database; The positioning module is configured to construct a regular expression according to the extreme words, and perform a full-text search on the book detail page according to the regular expression to locate the target paragraph containing the extreme words; The analysis module is configured to perform semantic feature analysis on the target paragraph through the preset model to determine whether the target extreme word in the target paragraph needs to be corrected; The execution module is configured to, when the target extreme word in the target paragraph needs to be corrected, generate a source file index link according to the file path of the target extreme word, and send the source file index link to a preset target.

[0014] 9 In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method described in any one of the above.

[0015] 10 In the fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described in any one of the above is executed.

[0016] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By obtaining the extreme words of books from the rule libraries of multiple e-commerce platforms at preset intervals and updating the local database accordingly, it is ensured that the extreme word library of the monitoring and management platform always remains up-to-date, can respond in a timely manner to the update of the definition of extreme words by e-commerce platforms, and thus effectively avoids the appearance of illegal content in the book detail pages. 2. Using the trained preset model to perform semantic feature analysis on the target paragraph can more accurately determine whether the target extreme word needs to be corrected, which is more intelligent and accurate than simply relying on keyword matching, reduces the possibility of false alarms and missed reports, and improves the processing efficiency. 3. By constructing a regular expression to perform a full-text search on the book detail page, the target paragraph containing extreme words can be quickly located. When it is determined that the target extreme word needs to be corrected, a source file index link is generated according to the file path of the target extreme word and sent to a preset target (such as the editor or manager of the book detail page), enabling the correction work to be carried out quickly and shortening the processing time. 4. Through automated and intelligent means, the extreme words in the book detail pages are monitored and managed, effectively improving the compliance of the book content, reducing the legal risks and economic losses caused by illegal content. By ensuring that no illegal extreme words appear in the book detail pages, the reading experience and trust of users can be improved, which helps to enhance the brand image and reputation of e-commerce platforms and book publishers. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic flowchart of the method for monitoring and managing extreme words of books disclosed in the embodiments of the present application; Figure 2 is a schematic block diagram of the system for monitoring and managing extreme words of books disclosed in the embodiments of the present application; Figure 3 is a schematic structural diagram of an electronic device disclosed in the embodiments of the present application.

[0018] Description of the reference numerals: 201, acquisition module; 202, positioning module; 203, analysis module; 204, execution module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments.

[0020] In the description of the embodiments of this application, words such as "for example" or "for illustration" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for illustration" is intended to present the relevant concepts in a specific manner.

[0021] In the description of the embodiments of this application, the term "a plurality of" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0022] This embodiment discloses a method for monitoring and managing book extreme words, which is applied to a monitoring and management platform. Figure 1 It is a schematic flow chart of the method for monitoring and managing book extreme words disclosed in the embodiments of this application, as Figure 1 shown. The method includes the following steps: S110: Obtain the extreme words of books from the rule libraries of multiple e-commerce platforms at preset intervals, update the local database according to the extreme words, and train a preset model according to the local database; Book extreme words refer to those words that exaggerate the content, quality, effect or influence of a book, and usually have the characteristics of absolutization and exaggeration. These words may include "the most", "first", "only", "absolute", etc., as well as words or phrases with similar meanings. E-commerce platforms usually have their own rule libraries to regulate the product information released by merchants, including restricting the use of certain extreme words. The system extracts extreme words related to book products from these platform rule libraries at preset time intervals (such as every day, every hour, etc.). The obtained extreme words are used to update a local database. This database may store information such as historically obtained extreme words and rule changes of each platform for subsequent analysis and processing. The update operation may include adding newly obtained extreme words, deleting no-longer-used extreme words, or updating the status of existing extreme words, etc. Using the updated local database as the data source, the system trains a preset model. This model may be used to identify extreme words in product descriptions, evaluate the usage frequency or impact of extreme words, etc. Training the model usually involves using a large amount of data (here it is extreme word data) to adjust the parameters of the model so that it can more accurately complete specific tasks. For example, if the model is used to identify extreme words, the training process may let the model learn how to distinguish extreme words from non-extreme words. The preset model is designed in advance and has specific structures and functions. Through training, the model can gradually adapt to new data and improve its performance on specific tasks.

[0023] Optionally, training the preset model according to the local database includes: Extracting a data set from the local database, the data set including extreme words and the context information of the extreme words; Cleaning the data set, converting the data in the cleaned data set into numerical feature vectors, and dividing the cleaned data set into a training set and a test set according to a preset ratio; Using a neural network architecture to train the preset model through the training set to learn the feature representation and context dependence of extreme words until a preset number of iterations or the loss function converges; Evaluating the preset model using the test set, and adjusting hyperparameters or the model structure according to the evaluation results until the model performance meets the requirements.

[0024] Extract a dataset containing extreme words and their context information from the updated local database. This dataset is the basis for subsequent model training and evaluation. The dataset should contain a rich sample of extreme words and their usage in contexts such as book descriptions and advertising copy. Clean the extracted dataset to remove noise data (such as invalid characters, duplicate data, etc.) to ensure data quality and consistency. The cleaning process may include steps such as data deduplication, missing value handling, and outlier detection. Convert the text data (such as extreme words and context information) in the cleaned dataset into numerical feature vectors. This typically involves text vectorization techniques such as the Bag of Words model, TF-IDF (Term Frequency-Inverse Document Frequency), word embeddings (such as Word2Vec, BERT), etc. Numerical feature vectors are the input form that machine learning models can process. Divide the cleaned dataset into a training set and a test set according to a preset ratio (such as 80% training set and 20% test set). The training set is used to train the model, and the test set is used to evaluate the performance of the model. Select a suitable neural network architecture (such as Convolutional Neural Network CNN, Recurrent Neural Network RNN, attention mechanism model, etc.) as the preset model. Use the training set to train the preset model through the backpropagation algorithm so that it learns the feature representations and context dependencies of extreme words. The training process includes multiple iterations, and in each iteration, the weights of the model are updated according to the gradient of the loss function. The training process continues until the preset number of iterations is reached or the loss function converges (i.e., the loss value no longer decreases significantly). Use the test set to evaluate the trained preset model and calculate the performance metrics of the model (such as accuracy, recall, F1 score, etc.). The evaluation results are used to measure the generalization ability of the model on unseen data. According to the evaluation results, adjust the hyperparameters (such as learning rate, batch size, number of network layers, etc.) or the model structure (such as adding new network layers, changing activation functions, etc.) of the model. Repeat the process of model training and evaluation until the model performance meets the requirements (i.e., reaches the preset performance metrics or no longer improves significantly).

[0025] Extracting a dataset containing extreme words and their contextual information from the local database helps the model understand the context and pattern of extreme words in actual use. The data cleaning process ensures the quality and consistency of the input data and avoids the impact of noisy data on model training. Converting the data into numerical feature vectors enables the model to process and understand the data. Using a neural network architecture for training can capture the feature representation and contextual dependencies of extreme words, improving the recognition accuracy of the model. By continuously adjusting hyperparameters (such as learning rate, batch size, etc.) and model structure (such as the number of layers, number of neurons, etc.), the performance of the model can be further optimized to make it more adaptable to different application scenarios and data distributions. This continuous optimization capability enables the model to continue to improve as the data is updated and changed, maintaining the advancement and practicality of the model. The trained model can be applied to the automatic review and monitoring of book content, helping publishers and e-commerce platforms to quickly identify and filter out book descriptions containing extreme words, and improve the compliance and quality of the content. This not only helps to maintain the good image and user experience of the platform, but also effectively avoids legal risks and disputes caused by improper use of extreme words.

[0026] S120, constructing a regular expression according to the extreme word, and performing a full-text search on the book detail page according to the regular expression to locate a target paragraph containing the extreme word; Design a regular expression based on the limit words. Regular expressions are a powerful text processing tool that can match and locate strings that meet specific patterns. For limit words, you can design a simple regular expression, such as \b最\b to match the single word "最" (where \b represents a word boundary, ensuring that the match is a complete word rather than a part of a word). For more complex limit words or phrases, you may need to design a more complex regular expression. Determine the book detail page text to be searched. This usually includes parts of the book such as the title, description, author introduction, and reviews. Apply the constructed regular expression to the target text for a full-text search. During this process, the algorithm traverses the entire text to find strings that match the regular expression. When the regular expression finds a match in the text, the position of the match (such as the start and end index) and the matched string itself are returned. Based on the position information of the match, you can determine the paragraph containing the limit word. This usually requires dividing the text into multiple paragraphs (if it has not been divided yet), and then checking each paragraph to see if it contains a match. Extract the target paragraphs containing the limit words, which may be paragraphs that need further review, modification, or marking.

[0027] Optionally, performing a full-text search on the book details page according to the regular expression to locate a target paragraph containing the extreme word includes: Traversing the text content in the book detail page based on the regular expression to match the selected paragraphs containing the extreme words; Convert the to-be-selected paragraph into a hash value, and deduplicate the to-be-selected paragraph according to the hash value to select a target paragraph from the to-be-selected paragraph.

[0028] Use a previously constructed regular expression (which can match extreme words in the book details page) to traverse the text content of the book details page. This traversal process is usually from the beginning to the end of the text, checking for matches word by word or character by character. Whenever the regular expression finds a match (i.e., an extreme word) in the text, the paragraph (or sentence, depending on the specific implementation) where it is located is regarded as a to-be-selected paragraph. A to-be-selected paragraph may contain multiple extreme word matches or only one. To remove duplicate to-be-selected paragraphs (for example, when the same paragraph is cited multiple times or appears repeatedly in different parts of the details page), each to-be-selected paragraph can be converted into a unique hash value. A hash function is an algorithm that converts data of any length into a hash value of a fixed length. It can map similar inputs to similar outputs (although not absolutely unique), but different inputs usually produce completely different hash values. For each to-be-selected paragraph, apply the hash function to generate a hash value. This hash value will be used as the unique identifier of the paragraph for subsequent deduplication steps. After generating the hash values of all to-be-selected paragraphs, compare the hash values. If the hash values of two or more to-be-selected paragraphs are the same, they are considered duplicates, and only one of them is retained as the target paragraph. After deduplication by hash value, the remaining to-be-selected paragraphs are the final target paragraphs. These paragraphs contain all the non-duplicate text parts containing extreme words in the book details page.

[0029] Through regular expressions, the text content of the book details page can be matched quickly and accurately to find the to-be-selected paragraphs containing extreme words. Regular expressions are a powerful text processing tool that can match strings that conform to specific patterns, ensuring the accuracy and integrity of the matching results. Converting the to-be-selected paragraph into a hash value is a process of converting data of any length into a string of a fixed length. Hash values are unique (ideally, different inputs produce different hash values), which makes hash values an effective tool for deduplication. Deduplicating the to-be-selected paragraphs based on hash values can effectively avoid the interference of duplicate paragraphs. In practical applications, there may be duplicate paragraphs or similar descriptions in the book details page. Deduplication by hash value can ensure that each target paragraph is unique, thereby improving the efficiency and accuracy of subsequent processing. By deduplication with hash values, the amount of data to be processed can be significantly reduced, the consumption of computing resources can be reduced, and the processing efficiency can be improved. After removing duplicate paragraphs, the target paragraphs are more refined and accurate, which helps with subsequent content review, modification, or marking work.

[0030] S130. Analyze the semantic features of the target paragraph through the preset model to determine whether the target extreme words in the target paragraph need to be corrected; The target paragraphs obtained after regular expression matching and hash value deduplication are input into the preset model. These paragraphs are text fragments containing potentially extreme words that need to be corrected. The preset model will perform in-depth analysis on the target paragraphs and extract their semantic features. These features may include semantic vectors of words, syntactic structures of sentences, topic information of paragraphs, etc. The model will also consider the context information in the target paragraph to more accurately understand the meaning and impact of extreme words in a specific context. This helps the model judge whether the use of extreme words is appropriate and whether correction is needed. Based on semantic feature analysis and context understanding, the preset model will evaluate the extreme words in the target paragraph. It will judge whether these extreme words are too exaggerated or absolute, whether they may mislead consumers or violate relevant laws and regulations. If the model believes that an extreme word needs to be corrected, it will generate corresponding correction suggestions. These suggestions may include replacing with more accurate words, deleting extreme words, or adjusting the expression of sentences, etc. The analysis results and correction suggestions of the model will be presented to relevant personnel (such as content reviewers, editors, etc.) in a certain form (such as text, marks, or visual interfaces). Relevant personnel can make corresponding modifications to the target paragraph according to the analysis results and correction suggestions of the model to ensure the compliance and accuracy of the book content. At the same time, these results can also be used as a reference for subsequent content review and monitoring.

[0031] Optionally, the analyzing the semantic features of the target paragraph through the preset model to determine whether the target extreme words in the target paragraph need to be corrected includes: Convert the text in the target paragraph into a semantic vector through the preset model, and determine the target feature vector of the target extreme word from the semantic vector; Match the usage of the target extreme word in the context with the standard usage in the preset rule library. When the usage is inconsistent with the standard usage, determine that the target extreme word needs to be corrected; When the usage is consistent with the standard usage, match the target feature vector with the standard semantic vector in the preset rule library. When the target feature vector is inconsistent with the standard semantic vector, determine that the target extreme word needs to be corrected.

[0032] Use a pre-trained preset model to convert the text in the target paragraph into semantic vectors. These semantic vectors are representations of the text in a specific semantic space, capable of capturing the semantic features and context relationships of the text. Among the converted semantic vectors, further determine the target feature vector of the target limit word. This is usually achieved by locating the corresponding part of the target limit word in the semantic vector, which may involve operations such as segmentation, weighting, or clustering of the semantic vector. Match the usage of the target limit word in the context with the standard usage in the preset rule base. The preset rule base may contain a series of norms and standards regarding the use of limit words, which are formulated based on laws and regulations, industry guidelines, or language habits, etc. By comparing the usage of the target limit word in the context with the standard usage in the preset rule base, determine whether they are consistent. If the usage is inconsistent, it indicates that there may be a problem with the use of the target limit word and further review is required. When the usage of the target limit word in the context is consistent with the standard usage in the preset rule base, it is also necessary to further match its target feature vector with the standard semantic vector in the preset rule base. The standard semantic vector is the semantic representation of the standard usage of the limit word in the preset rule base. By comparing the target feature vector with the standard semantic vector, determine whether they are consistent. If the target feature vector is inconsistent with the standard semantic vector, it indicates that although the target limit word conforms to the norms in terms of usage, there may be deviations or misunderstandings in semantics, so it also needs to be corrected. Based on the above analysis, when the usage or semantics of the target limit word in the context is inconsistent with the standard in the preset rule base, determine that the target limit word needs to be corrected. The correction suggestions may include replacing it with a more accurate word, adjusting the sentence expression, or deleting the limit word, etc. Output the correction decision and the corresponding correction suggestions to relevant personnel (such as content reviewers, editors, etc.) so that they can make corresponding modifications to the target paragraph according to the suggestions.

[0033] Suppose there is a book called "Extreme Adventure". In this book, the author describes an adventure experience and uses the extreme word "absolutely". The original paragraph might be as follows: This adventure is absolutely the most thrilling one in my life. I have never seen such magnificent scenery nor felt such intense fear. Apply the above paragraph directly to the product details page. Use a trained natural language processing model (such as BERT) to convert the above original paragraph into semantic vectors. These semantic vectors can capture the word "absolutely" in the text and the related context information. In the obtained semantic vectors, locate the target feature vector corresponding to the word "absolutely". This may involve further processing of the semantic vectors, such as segmentation, weighting, etc., to highlight the features of the word "absolutely". Suppose the standard usage of the extreme word "absolutely" in the preset rule base is "Avoid using overly absolute words in objective descriptions to avoid misleading readers or causing ambiguity". Match the usage of the word "absolutely" in the original paragraph with the standard usage in the preset rule base. In this example, the word "absolutely" is used to describe the author's subjective feelings. Although it is somewhat subjective, it does not directly mislead readers or cause ambiguity. However, from the perspective of objective description, using the word "absolutely" may still seem overly absolute. Suppose the standard semantic vector of the extreme word "absolutely" in the preset rule base represents "a strong affirmation or assertion, but may lack objective basis or be misleading". Match the target feature vector of the word "absolutely" in the original paragraph with the standard semantic vector. In this example, although the word "absolutely" semantically expresses the author's strong feelings, there is still a certain deviation between it and the standard semantic vector. Especially when the word "absolutely" is used to describe objective facts or experiences, this deviation may be more obvious. Based on the above analysis, it can be considered that although the word "absolutely" in the original paragraph expresses the author's subjective feelings to a certain extent, from the perspectives of objectivity and accuracy, its usage may not be appropriate. However, the standard semantic vector also stipulates that no correction is required if it is a quoted sentence. Therefore, the word "absolutely" can be left uncorrected.

[0034] By converting the text in the target paragraph into semantic vectors and determining the target feature vector of the target limit word, this process can deeply mine the semantic information in the text, thereby more accurately judging the specific meaning and usage of the target limit word in the context. This semantic-level analysis is more accurate than traditional text matching methods and can capture more subtle semantic differences. The standard usages and standard semantic vectors in the preset rule library provide norms and standards for the use of limit words. By matching the usage of the target limit word in the context with the standard usages in the preset rule library and matching the target feature vector with the standard semantic vector, it can flexibly handle the usage of limit words in different contexts. This flexibility enables this process to be applicable to different types of books and text contents, improving its versatility and practicality.

[0035] Optionally, the method further includes: Receiving the correction result of the preset target for the target limit word, where the correction result includes rejecting the correction of the first target limit word and the content after correcting the second target limit word; Performing negative training on the preset model according to the first target limit word and the context information of the first target limit word to reduce the recognition of the first target limit word by the preset model; Performing positive training on the preset model according to the second target limit word and the content after correcting the second target limit word to increase the recognition of the second target limit word by the preset model.

[0036] The system has received the correction results for the target extreme words. These correction results may come from human reviewers, automated review systems, or other sources. The correction results generally include two parts: one part is the first target extreme words that are rejected for correction, that is, those extreme words that are considered not to need modification after review; the other part is the second target extreme words after correction and their corresponding content, that is, those extreme words that are considered to need modification after review and have been modified, along with their new content. The purpose of negative training is to reduce the misrecognition of the first target extreme words by the preset model. That is to say, if an extreme word is used appropriately in the context without causing ambiguity or misleading readers, then unnecessary correction of that word should be avoided. Through negative training, the model can learn the situations where the use of extreme words is appropriate, thus reducing the misjudgment of them. When implementing negative training, the system will use the first target extreme words and their context information as training data. These data will be input into the preset model, and the parameters of the model will be adjusted to reduce the recognition rate of this type of extreme words by the model. In this way, when the model encounters a similar situation again, it is more likely to correctly judge whether the use of the extreme word is appropriate. The purpose of positive training is to increase the recognition ability of the preset model for the second target extreme words. That is to say, if an extreme word is used inappropriately in the context and needs to be modified, then it should be ensured that the model can accurately identify this and give corresponding correction suggestions. Through positive training, the model can learn the situations where the use of extreme words is inappropriate and improve its recognition accuracy. When implementing positive training, the system will use the second target extreme words and their corrected content as training data. These data will also be input into the preset model, and the parameters of the model will be adjusted to increase the recognition rate of this type of extreme words by the model. In this way, when the model encounters a similar situation again, it can more accurately judge whether the use of the extreme word is inappropriate and give correct correction suggestions.

[0037] When the result of rejecting the correction of the first target extreme word is received, it indicates that the preset model may misjudge the recognition of this extreme word. By performing negative training on the preset model, that is, reducing the model's recognition of this extreme word, the misjudgment rate in future similar situations can be reduced, and the recognition accuracy of the model can be improved. When the content after correcting the second target extreme word is received, it indicates that the preset model's recognition of this extreme word is accurate, but it may need to be further optimized to better adapt to different contexts. By performing positive training on the preset model, that is, increasing the model's recognition of this extreme word, the model's recognition ability in future similar contexts can be enhanced. By receiving the correction result and training the preset model, an iterative update process is formed. This process enables the model to continuously learn from actual data and optimize its own recognition ability, so as to better meet the extreme word recognition requirements in different text contents and contexts. As the model is continuously trained and optimized, its adaptability will gradually increase. This means that the model can more accurately recognize extreme words in different types and contexts, thus improving its value in practical applications.

[0038] Optionally, the negative training of the preset model according to the first target extreme word and the context information of the first target extreme word to reduce the recognition of the first target extreme word by the preset model includes: Adding the first target extreme word and the context information of the first target extreme word as negative samples to the training set; Using the updated training set to retrain the preset model to reduce the recognition rate of the negative samples by the preset model.

[0039] The first target extreme word refers to an extreme word that is wrongly marked by the model as needing correction in the initial semantic feature analysis, but is actually used appropriately in the context and does not need correction. Context information refers to the sentence or paragraph containing the first target extreme word, which provides the necessary background for understanding the use of the extreme word in a specific context. In machine learning, negative samples usually refer to those data points that do not belong to the target class. In this scenario, negative samples refer to those first extreme words wrongly marked by the model as needing correction and their context information. Adding the first target extreme word and its context information as negative samples to the training set means adding more information about "when extreme words should not be corrected" to the training set. This helps the model to more accurately judge the extreme words that really need correction and those that do not need correction when encountering similar situations in the future. Retraining the preset model with the updated training set to reduce the recognition rate of the preset model for negative samples. The updated training set refers to the training set after adding negative samples. It now contains more information about the use of extreme words in different contexts, including which ones need correction and which ones do not. Retraining the preset model refers to the process of retraining the preset model with the updated training set. Through this process, the model can learn more rules about the use of extreme words in different contexts, thus reducing the misrecognition of negative samples. The purpose of retraining the preset model is to enable the model to better adapt to the new data, that is, those extreme words wrongly marked as needing correction but actually not needing correction and their context information. Through this process, the model can gradually reduce the misrecognition of such extreme words and improve the recognition accuracy.

[0040] By adding negative samples containing the first target extreme word and its context information to the training set and retraining the model, the model can learn to better distinguish between texts containing inappropriate or extreme expressions and normal or expected texts. This helps to reduce the situation of model false alarms (i.e., misjudging normal texts as containing extreme words), thus improving the accuracy and overall performance of the model. It allows adjusting the sensitivity of the model to specific types of vocabulary according to actual needs. For example, in the review of advertising copy, it may be necessary to strictly restrict the use of certain exaggerated or misleading extreme words, and through this method, the recognition ability of the model for these words can be flexibly adjusted. By introducing more diverse negative samples, it can help the model learn a wider range of context information and vocabulary usage patterns, thus reducing the overfitting of the model to a specific data set and reducing the model bias caused by training data bias. In scenarios where strict control of text quality is required (such as social media, online advertising platforms, etc.), reducing the misrecognition of extreme words by the model can improve the user experience and avoid user dissatisfaction or complaints caused by false alarms.

[0041] S140. When the target extreme word in the target paragraph needs to be corrected, generate a source file index link according to the file path of the target extreme word, and send the source file index link to a preset target.

[0042] At a certain stage of content review or text processing, the system has performed semantic feature analysis on the target paragraph through a preset model and determined whether there are target extreme words that need to be corrected. If it is detected that the use of the target extreme word does not conform to the preset rules or standards, the system will mark these words as needing correction. For each target extreme word that needs to be corrected, the system will generate a unique index link based on its position information in the source file (such as file path, paragraph number, line number, etc.). This index link is actually a pointer to a specific location in the source file, which allows the recipient to quickly locate the text paragraph containing the target extreme word. The system will send the generated source file index link to a preset target object. This target object may be a staff member responsible for content editing or review, or another automated system for further processing or marking these content that needs to be corrected. The sending method can be email, instant message, API call, or any other preset communication means. After receiving the source file index link, the preset target can quickly locate the corresponding position in the source file according to the link, view and understand why the target extreme word needs to be corrected. The preset target can perform corresponding correction operations, such as replacing with more appropriate words, adjusting the sentence structure, or deleting unnecessary expressions. After the correction is completed, if the system supports it, the preset target can also feedback the correction result to the system for further verification or updating of the preset model.

[0043] Optionally, the generating a source file index link according to the file path of the target extreme word includes: Parse the file path where the target extreme word is located, and extract key information, where the key information includes file name, file type, and storage location; Construct a source file index database according to the key information, where the source file index database contains a unique identifier, path information, and modification time of the source file; Generate a corresponding source file index link according to the unique identifier.

[0044] The system needs to parse the file path where the target limit word is located. This path is usually a string that contains the complete path from the root directory to the target file. During the process of parsing the path, the system extracts key information, which usually includes the file name (i.e., the name of the target file), the file type (such as.txt,.docx,.pdf, etc., indicating the format or type of the file), and the storage location (i.e., the specific location of the file in the file system, which may be the path of a folder). Based on the extracted key information, the system designs a source file index database. This database is used to store metadata such as the unique identifier of the source file, path information, and modification time. Each source file will have a unique identifier in the database (such as UUID, file hash value, etc.) to ensure the uniqueness of each file in the database. The path information includes the complete path or relative path of the file to locate the position of the source file in the file system. The modification time records the timestamp when the source file was last modified, which helps to track the update situation of the file. After extracting the key information and designing the database, the system enters this information into the source file index database to form a complete source file record. With the source file index database, the system can generate the corresponding source file index link according to the unique identifier of the source file. This link is usually a URL or a similar identifier that contains enough information to locate the position of the source file in the database. The format of the link can be designed according to actual needs, but it usually should contain some key elements, such as the address of the database server, the name of the database, the unique identifier of the source file, etc. When generating the link, security issues also need to be considered. For example, encryption or hashing processing can be used to ensure the uniqueness and security of the link and prevent unauthorized access.

[0045] By parsing the file path where the target extreme word is located, the system can extract key information such as the file name, file type, and storage location. This information serves as the basis for building the source file index database. The establishment of the source file index database enables the system to store and manage relevant information of source files in a structured manner. This significantly improves the efficiency of information retrieval because the system can quickly locate specific source files. Generating corresponding source file index links based on the unique identifiers of source files in the index database allows users or the system to directly access the source files through the links without having to manually search or enter complex file paths. Each source file has a unique identifier in the index database, which ensures the uniqueness and traceability of the source files. Even if the file name or storage location changes, as long as the unique identifier remains the same, the system can still accurately locate the source file. The source file index database also contains the path information and modification time of the source files. This information is very important for version control and data recovery. For example, when it is necessary to roll back to a specific version of a source file, the system can quickly find the corresponding version based on the modification time. Through the source file index database, the system can also implement more fine-grained permission management. For example, certain users or systems can be restricted to accessing only specific types of files or files within a specific time period.

[0046] This embodiment also discloses a monitoring and management system for book extreme words. Figure 2 It is a schematic diagram of the modules of the monitoring and management system for book extreme words disclosed in the embodiments of the present application, as Figure 2 shown. The system includes a collection module 201, a positioning module 202, an analysis module 203, and an execution module 204, where: The collection module 201 is configured to obtain extreme words of books from the rule libraries of multiple e-commerce platforms at preset intervals, update the local database according to the extreme words, and train a preset model according to the local database; The positioning module 202 is configured to construct a regular expression according to the extreme words and perform a full-text search on the book detail page according to the regular expression to locate the target paragraph containing the extreme words; The analysis module 203 is configured to perform semantic feature analysis on the target paragraph through the preset model to determine whether the target extreme words in the target paragraph need to be corrected; The execution module 204 is configured to generate a source file index link according to the file path of the target extreme word when the target extreme words in the target paragraph need to be corrected, and send the source file index link to a preset target.

[0047] Optionally, the collection module 201 is configured to: Extract a data set from the local database, where the data set includes extreme words and context information of the extreme words; Clean the data set, convert the data in the cleaned data set into numerical feature vectors, and divide the cleaned data set into a training set and a test set according to a preset ratio; Use a neural network architecture to train the preset model through the training set to learn the feature representation and context dependence of extreme words until a preset number of iterations is reached or the loss function converges; Use the test set to evaluate the preset model, and adjust hyperparameters or the model structure according to the evaluation results until the model performance meets the requirements.

[0048] Optionally, the positioning module 202 is configured to: Traverse the text content in the book details page based on the regular expression to match candidate paragraphs containing extreme words; Convert the candidate paragraphs into hash values, and remove duplicates from the candidate paragraphs according to the hash values to select target paragraphs from the candidate paragraphs.

[0049] Optionally, the analysis module 203 is configured to: Convert the text in the target paragraph into a semantic vector through the preset model, and determine the target feature vector of the target extreme word from the semantic vector; Match the usage of the target extreme word in the context with the standard usage in the preset rule library. When the usage is inconsistent with the standard usage, determine that the target extreme word needs to be corrected; When the usage is consistent with the standard usage, match the target feature vector with the standard semantic vector in the preset rule library. When the target feature vector is inconsistent with the standard semantic vector, determine that the target extreme word needs to be corrected.

[0050] Optionally, the system further includes an adjustment module, and the adjustment module is configured to: Receive the correction result of the preset target for the target extreme word, where the correction result includes rejecting the correction of the first target extreme word and the content after correcting the second target extreme word; Perform negative training on the preset model according to the first target extreme word and the context information of the first target extreme word to reduce the recognition of the first target extreme word by the preset model; Perform positive training on the preset model according to the second target extreme word and the content after correcting the second target extreme word to increase the recognition of the second target extreme word by the preset model.

[0051] Optionally, the adjustment module is configured to: Add the first target extreme word and the context information of the first target extreme word to the training set as negative samples; Retrain the preset model using the updated training set to reduce the recognition rate of the preset model for negative samples.

[0052] Optionally, the execution module 204 is configured to: Parse the file path where the target extreme word is located, and extract key information, where the key information includes the file name, file type, and storage location; Construct a source file index database according to the key information, where the source file index database includes the unique identifier, path information, and modification time of the source file; Generate a corresponding source file index link according to the unique identifier.

[0053] It should be noted that when the device provided in the above embodiment realizes its functions, only the division of the above function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.

[0054] This embodiment also discloses an electronic device. Refer to Figure 3 , the electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, a network interface 304, and at least one memory 305.

[0055] Among them, the communication bus 302 is used to realize the connection and communication between these components.

[0056] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0057] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0058] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server using various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by invoking the data stored in the memory 305, it performs various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.

[0059] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. As Figure 3 shown, the memory 305, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for the monitoring and management method of book extreme words.

[0060] In Figure 3In the electronic device shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; while the processor 301 can be used to call the application program for monitoring and managing the book limit words stored in the memory 305. When executed by one or more processors 301, the electronic device is caused to execute the method of one or more of the above embodiments.

[0061] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0062] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0063] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0064] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0065] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0066] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. And the aforementioned memory 305 includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0067] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the disclosure of the specification, those skilled in the art will readily think of other implementation schemes of the present disclosure. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A monitoring and management method for book limit words, characterized in that: Applied to a monitoring management platform, the method includes: Obtaining the extreme words of books from the rule bases of multiple e-commerce platforms at preset intervals, updating a local database according to the extreme words, and training a preset model according to the local database; Constructing a regular expression according to the extreme word, and performing a full-text search on the book detail page according to the regular expression to locate a target paragraph containing the extreme word; Performing semantic feature analysis on the target paragraph using the preset model to determine whether the target limit words in the target paragraph need to be modified; When the target limit word in the target paragraph needs to be revised, a source file index link is generated according to the file path of the target limit word, and the source file index link is sent to a preset target.

2. The monitoring and management method of book limit words according to claim 1 is characterized in that: The training of the preset model according to the local database comprises: extracting a data set from the local database, the data set including extreme words and context information of the extreme words; Cleaning the data set, converting the data in the cleaned data set into numerical feature vectors, and dividing the cleaned data set into a training set and a test set according to a preset ratio; Using a neural network architecture to train the preset model through the training set to learn the feature representation and context dependency of the extreme words until a preset number of iterations is reached or the loss function converges; The preset model is evaluated using the test set, and the hyperparameters or model structure are adjusted according to the evaluation results until the model performance meets the requirements.

3. The monitoring and management method of book limit words according to claim 1 is characterized in that: The full-text search of the book details page according to the regular expression to locate the target paragraph containing the extreme word includes: Traversing the text content in the book details page based on the regular expression to match the selected paragraphs containing the extreme words; The candidate paragraphs are converted into hash values, and the candidate paragraphs are deduplicated according to the hash values ​​to select a target paragraph from the candidate paragraphs.

4. The monitoring and management method of book limit words according to claim 1 is characterized in that: The performing semantic feature analysis on the target paragraph by using the preset model to determine whether the target limit words in the target paragraph need to be modified includes: Converting the text in the target paragraph into a semantic vector by using the preset model, and determining a target feature vector of the target limit word from the semantic vector; Matching the usage of the target limit word in the context with the standard usage in a preset rule base, and when the usage is inconsistent with the standard usage, determining that the target limit word needs to be revised; When the usage is consistent with the standard usage, the target feature vector is matched with the standard semantic vector in the preset rule base; when the target feature vector is inconsistent with the standard semantic vector, it is determined that the target limit word needs to be modified.

5. The monitoring and management method of book limit words according to claim 1 is characterized in that: The method further comprises: receiving a modification result of the preset target on the target limit word, wherein the modification result includes refusing to modify the first target limit word and modifying the content after the second target limit word; Performing negative training on the preset model according to the first target limit word and context information of the first target limit word to reduce recognition of the first target limit word by the preset model; The preset model is forward trained according to the second target limit word and the content after the modified second target limit word to increase the recognition of the second target limit word by the preset model.

6. The monitoring and management method of book limit words according to claim 5 is characterized in that: The negative training of the preset model according to the first target limit word and the context information of the first target limit word to reduce the recognition of the first target limit word by the preset model includes: adding the first target limit word and the context information of the first target limit word as negative samples to a training set; The preset model is retrained using the updated training set to reduce the recognition rate of the preset model for negative samples.

7. The monitoring and management method of book limit words according to claim 1 is characterized in that: Generating a source file index link according to the file path of the target limit word comprises: Parsing the file path where the target limit word is located, and extracting key information, wherein the key information includes the file name, file type and storage location; Building a source file index database according to the key information, the source file index database including a unique identifier, path information and modification time of the source file; A corresponding source file index link is generated according to the unique identifier.

8. A monitoring and management system for book limit words, characterized in that: It includes acquisition module, positioning module, analysis module and execution module, among which: A collection module, configured to obtain the extreme words of books from the rule bases of multiple e-commerce platforms at preset intervals, update the local database according to the extreme words, and train a preset model according to the local database; A positioning module, configured to construct a regular expression according to the extreme word, and perform a full-text search on the book detail page according to the regular expression to locate a target paragraph containing the extreme word; An analysis module configured to perform semantic feature analysis on the target paragraph through the preset model to determine whether the target limit words in the target paragraph need to be modified; The execution module is configured to generate a source file index link according to the file path of the target limit word when the target limit word in the target paragraph needs to be modified, and send the source file index link to a preset target.

9. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.