Text content analysis method and system, computer device, medium, product
By performing content parsing and multi-dimensional analysis on the text, and calling upon the analytical agent to obtain similarity and confidence scores, the unscientific and inefficient problems of traditional article review methods are solved, achieving efficient and accurate evaluation of text quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING QINGSONG YIKANG INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-09-28
- Publication Date
- 2026-04-17
AI Technical Summary
Traditional article review methods lack scientific basis, resulting in incomplete reviews or wasted resources. The scoring rules are inconsistent, highly subjective, difficult to quantify, and inefficient.
By parsing the text to be analyzed, extracting keywords, technical fields, and complexity, determining the text level, calling the corresponding number of analytical agents, obtaining similarity and confidence scores, and integrating multi-dimensional analysis results, an objective and accurate text quality assessment can be achieved.
It enables a comprehensive and objective evaluation of text content, improves evaluation efficiency and accuracy, adapts to different needs and scenarios, allocates resources reasonably, avoids redundant analysis, and provides clear quality evaluation results.
Smart Images

Figure CN121328519B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of text analysis technology, and in particular to a text content analysis method and system, computer device, medium, and product. Background Technology
[0002] In the modern information society, articles serve as a crucial carrier of information dissemination, and their quality is vital to the effectiveness and value realization of communication. Different types of articles, such as academic papers, technical documents, and business reports, all require accurate quality assessments to ensure they play their due role in their respective fields. Traditionally, article quality assessment relied primarily on expert commentary, where experts, with their professional knowledge and experience, could conduct in-depth analyses of articles. However, this approach is no longer adequate for modern needs.
[0003] Traditional article review methods have a series of problems. First, the number of expert reviews lacks scientific basis, leading to incomplete reviews or wasted resources. Second, there is no clear limit to the number of reviews per article, which easily leads to repetitive comments and low efficiency. Third, the scoring rules are not uniform, the review results are highly subjective, and it is difficult to quantify the quality of the articles. Fourth, manual review and scoring calculation are time-consuming, affecting the efficiency and accuracy of the evaluation. Summary of the Invention
[0004] In view of this, the present disclosure provides a text content analysis method and system, computer device, medium, and product that can solve the problems of poor accuracy and low efficiency of text evaluation results in the prior art.
[0005] In a first aspect, embodiments of this disclosure provide a text content analysis method, including:
[0006] The text to be analyzed is parsed to extract target information, which includes keywords, technical fields, and text complexity.
[0007] Determine the text level based on the target information;
[0008] The corresponding number of analytical agents are invoked based on the text level;
[0009] Based on each of the aforementioned analytical agents, content quality analysis information, language expression analysis information, and structural arrangement analysis information are obtained;
[0010] Obtain the similarity between the target information and each of the analytical agents, and obtain the confidence level of each analytical agent based on the similarity.
[0011] Based on the content quality analysis information, language expression analysis information, structural arrangement analysis information, and confidence level corresponding to all the analytical agents, the content quality dimension analysis results, language expression dimension analysis results, and structural arrangement dimension analysis results are obtained.
[0012] Based on the analysis results of the content quality dimension, the language expression dimension, and the structure and arrangement dimension, the text content analysis results are obtained.
[0013] Secondly, embodiments of this disclosure also provide a text content analysis system, including:
[0014] The information extraction module is used to parse the text to be analyzed and extract target information, including keywords, technical fields, and text complexity.
[0015] A text level determination module is used to determine the text level based on the target information;
[0016] The agent invocation module is used to invoke a corresponding number of analytical agents according to the text level; and to obtain content quality analysis information, language expression analysis information, and structural arrangement analysis information based on each analytical agent.
[0017] The confidence acquisition module is used to acquire the similarity between the target information and each of the analysis agents, and to acquire the confidence of each analysis agent based on the similarity.
[0018] The multi-dimensional analysis module is used to obtain content quality dimension analysis results, language expression dimension analysis results, and structural arrangement dimension analysis results based on the content quality analysis information, language expression analysis information, structural arrangement analysis information, and confidence scores corresponding to all the analysis agents.
[0019] The content analysis result acquisition module is used to obtain text content analysis results based on the content quality dimension analysis results, the language expression dimension analysis results, and the structure and arrangement dimension analysis results.
[0020] Thirdly, this disclosure also provides a computer device, which adopts the following technical solution:
[0021] The computer device includes:
[0022] At least one processor; and,
[0023] A memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform any of the text content analysis methods described above.
[0025] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to perform any of the text content analysis methods described above.
[0026] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.
[0027] The text content analysis method disclosed in this application first parses the text to be analyzed, extracting target information, including keywords, technical fields, and text complexity. Then, it determines the text level based on the target information and calls upon a corresponding number of analysis agents based on the text level to ensure the accuracy and comprehensiveness of the analysis. Next, it obtains content quality analysis information, language expression analysis information, and structural arrangement analysis information based on each analysis agent, enabling a comprehensive evaluation of the text. Then, it obtains the similarity between the target information and each analysis agent, and obtains the confidence level of each analysis agent based on the similarity. Finally, it calculates the content quality score corresponding to all analysis agents. By analyzing information, language expression, and structural arrangement, along with confidence levels, we obtain analysis results for the content quality, language expression, and structural arrangement dimensions. Integrating the analytical information from multiple agents allows us to fully utilize the strengths of each agent while considering their credibility, resulting in more objective and accurate analysis results across all dimensions. Finally, based on these analysis results, we obtain the text content analysis results. By combining these results across all dimensions, we can provide a comprehensive and objective evaluation of the text content, offering users a clear text quality assessment.
[0028] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0029] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 This is a flowchart illustrating the text content analysis method provided in an embodiment of the present disclosure.
[0031] Figure 2 A flowchart illustrating the method for obtaining text complexity provided in this embodiment of the disclosure.
[0032] Figure 3 This is a flowchart illustrating the method for obtaining the confidence level of each analytical agent provided in embodiments of this disclosure.
[0033] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation
[0034] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0035] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0036] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.
[0037] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0038] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0039] Reference Figure 1 This application discloses a text content analysis method, including:
[0040] S100: Perform content parsing on the text to be analyzed to extract target information, including keywords, technical fields, and text complexity.
[0041] S200, determine the text level based on the target information;
[0042] S300 invokes the corresponding number of analytical agents based on the text level;
[0043] Each analytical agent obtains content quality analysis information, language expression analysis information, and structural arrangement analysis information.
[0044] S400: Obtain the similarity between the target information and each analytical agent, and obtain the confidence level of each analytical agent based on the similarity.
[0045] S500 obtains the analysis results for the content quality dimension, language expression dimension, and structure arrangement dimension based on the content quality analysis information, language expression analysis information, structure arrangement analysis information, and confidence scores of all analytical agents.
[0046] Specifically, the content quality dimension analysis result is N: .
[0047] The result of the language expression dimension analysis is Y: .
[0048] The structural arrangement dimension analysis results are as follows : .
[0049] in, For the first Content quality analysis information corresponding to each analytical agent. For the first The confidence level of each analytical agent, where P is the total number of analytical agents. For the first The language expression analysis information corresponding to each analytical agent. For the first The structural arrangement analysis information corresponding to each analytical agent.
[0050] S600 obtains text content analysis results based on the analysis results of content quality dimension, language expression dimension, and structural arrangement dimension.
[0051] The text content analysis results include at least one of the following: the overall analysis score of the text content, analysis information for each dimension, and corresponding improvement suggestions.
[0052] The text content analysis method disclosed in this application analyzes text from multiple perspectives, including content quality, language expression, and structural arrangement, enabling a comprehensive assessment of text quality. Through steps such as extracting target information, determining text levels, invoking an appropriate number of analytical agents, and calculating confidence scores, it fully considers the characteristics of the text and the credibility of the analytical agents, improving the accuracy of the analysis results. Furthermore, it can adjust parameters such as text level classification rules, the number of analytical agents invoked, and the weights of each dimension according to different needs and scenarios, exhibiting strong flexibility and adaptability. It rationally allocates analytical resources based on text levels, effectively avoiding the application of the same analysis methods and resource inputs to all texts, thus improving analysis efficiency. The text content analysis method disclosed in this application is an automated and standardized text analysis method that can effectively meet the requirements of modern information society for text evaluation, achieving high-efficiency and high-accuracy text assessment.
[0053] Reference Figure 2 Methods for obtaining text complexity include:
[0054] A100 obtains the number of words, technical terms, total words, total number of sentences, and the number of words, compound sentences, and cited references for each sentence based on the text to be analyzed.
[0055] Each sentence is marked by a period, exclamation mark, or question mark to identify it as a sentence.
[0056] A200, based on the preset word count range scoring rules, obtains the word count dimension score corresponding to the word count of the article.
[0057] Specifically, when the word count is between 0 and 500, the word count score is 1; when the word count is between 501 and 1500, the word count score is 2; and when the word count is above 1501, the word count score is 3. This scoring method allows the impact of word count on text complexity to be presented in a quantitative form.
[0058] A300 obtains the terminology density based on the number of professional terms and the total number of words, and obtains the terminology density dimension score corresponding to the professional terminology density according to the preset terminology density interval scoring rules.
[0059] The terminology density is calculated as the ratio of the number of specialized terms to the total number of words. A terminology density score of 1 is given when the density is between 0 and 0.1; 2 when it's between 0.11 and 0.3; and 3 when it's above 0.31. Calculating terminology density and assigning scores based on preset ranges objectively measures the professionalism of the text. Terminology density reflects the proportion of specialized vocabulary in the text, and different density ranges correspond to different scores, helping to distinguish texts of varying levels of professionalism. For example, a popular science article might have a low terminology density and a low score, indicating it's suitable for general readers; while a professional academic paper might have a high terminology density and a high score, indicating it's primarily aimed at professionals in the field.
[0060] A400 identifies all sentences whose word count exceeds the preset character limit and records them as long sentences; it obtains sentence density based on the number of compound sentences, the total number of long sentences, and the total number of sentences; and it obtains the sentence density dimension score corresponding to the sentence density according to the preset sentence density interval scoring rules.
[0061] Sentence density is the ratio of the sum of the number of complex sentences and the total number of long sentences to the total number of sentences. A sentence density score of 1 is given when the score is between 0 and 0.2; 2 when it is between 0.21 and 0.5; and 3 when it is above 0.51. Sentence density comprehensively considers the relationship between the number of complex sentences, the number of long sentences, and the total number of sentences, thus reflecting the overall complexity of the text's sentence structure. A higher sentence density indicates a greater proportion of complex sentences in the text, making it more difficult to understand.
[0062] A500 obtains the citation dimension score corresponding to the number of cited documents based on the preset citation interval scoring rules.
[0063] Specifically, when the number of cited references is 0, the score for the citation dimension is 1; when the number of cited references is between 1 and 3, the score is 2; and when the number of cited references is 4 or more, the score is 3. A higher number of cited references indicates that the text involves more specialized knowledge and previous research findings, requiring readers to have a certain level of professional background and knowledge to better understand the text.
[0064] A600 calculates the text complexity by weighting and summing the scores for word count, term density, sentence density, and citations according to preset weights.
[0065] By weighting and summing the scores of each dimension according to preset weights, the influence of multiple factors on text complexity can be comprehensively considered, resulting in a comprehensive and objective text complexity assessment. Different dimensions may have different degrees of importance in text complexity, and this difference can be reflected by setting weights.
[0066] The method for S200 "determining text levels based on target information" includes:
[0067] S210, the number of keywords and their coverage determine the text's professional depth score.
[0068] Specifically, the number of keywords in the text is counted. If the number is 0-3, a score of 1 is awarded; if it's 4-6, 2 points; if it's 7-9, 3 points; and if it's 10 or more, 4 points. If all keywords cover only a single, broad basic professional field, and the keywords are mostly basic concepts, the coverage level is classified as elementary, and the corresponding score is 1 point. If all keywords cover multiple aspects of a professional field, or cover a small number of related professional fields, the coverage level is classified as intermediate, and the corresponding score is 2 points. If all keywords cover multiple different professional fields, and the keywords cover professional concepts at different levels, the coverage level is classified as advanced, and the corresponding score is 3 points. If all keywords cover multiple professional fields, and have in-depth and comprehensive coverage of each field, the coverage level is classified as top-level, and the corresponding score is 4 points.
[0069] Finally, the quantity score and range score are weighted and summed according to preset weights to obtain the text professional depth score. This method can calculate the professional depth score of the text in a relatively objective and accurate manner, providing an important basis for subsequent text level assessment.
[0070] S220, determine the text difficulty score based on the technical field.
[0071] Specifically, this includes: determining the popularity and interdisciplinary nature of the technical field; and determining the text difficulty score based on the popularity and interdisciplinary nature of the technical field. Specifically, if the technical field is one that is generally known to the public (such as common sense, basic science, etc.) and does not involve more than two different technical fields, the text difficulty score is 1 point; if the technical field is one that is generally known to the public and involves more than two different technical fields, the text difficulty score is 2 points; if the technical field is not one that is generally known to the public and does not involve more than two different technical fields, the text difficulty score is 3 points; and if the technical field is not one that is generally known to the public and involves more than two different technical fields, the text difficulty score is 4 points.
[0072] In this step, the text difficulty score is determined by combining the popularity of the technical field and its interdisciplinary nature, which comprehensively considers the characteristics of the fields involved in the text. The popularity of the field determines the public's familiarity with it, while the interdisciplinary nature reflects the comprehensiveness and complexity of the text's knowledge. By using clear scoring rules, the text difficulty of different fields and interdisciplinary situations is quantified, avoiding the bias of subjective judgment.
[0073] S230 calculates the target comprehensive score by weighting the text's professional depth score, text difficulty score, and text complexity score according to preset weights.
[0074] By weighting and summing the text's professional depth score, text difficulty score, and text complexity score according to preset weights, the overall characteristics of the text can be comprehensively considered. Different factors have different importance in the text level assessment, and this difference can be accurately reflected by setting weights. By integrating the scores of multiple dimensions into a comprehensive score, a unified and objective assessment standard is provided for determining the text level, avoiding the one-sidedness of single-factor assessment and making the assessment results more accurate and reliable.
[0075] S240. The text level is determined based on the overall target score. The text level can be any one of the following: normal, medium, or complex.
[0076] Specifically, the correspondence between the overall score and the text level for different intervals can be flexibly set according to needs.
[0077] The method of "calling the corresponding number of analysis agents according to the text level" in S300 specifically includes: when the text level is normal, the text is sent to M analysis agents; when the text level is medium, the text is sent to N analysis agents; when the text level is complex, the text is sent to P analysis agents; where P > N > M ≥ 1.
[0078] More preferably, 1≤M≤2, 3≤N≤5, 6≤P≤10.
[0079] In this embodiment, different levels of text have different requirements for analysis and processing. Ordinary-level text is relatively simple, less specialized, and less complex, requiring only a small number of analytical agents to complete the analysis task, avoiding excessive resource investment. Complex-level text contains a large amount of specialized knowledge, complex logic, and rich details, requiring more computing resources and analytical capabilities. Calling on P analytical agents can fully leverage the collaborative effects of multiple agents to conduct in-depth analysis of the text from different perspectives, ensuring the accuracy and comprehensiveness of the analysis results. For example, in a cutting-edge academic paper involving multiple interdisciplinary fields, multiple agents can process content from different professional fields, jointly completing the interpretation of the paper.
[0080] As the text level increases, its information content and complexity also increase. Increasing the number of analytical agents ensures comprehensive coverage of the text and avoids missing important information. Multiple agents can analyze the text from different perspectives, complementing and verifying each other, thus improving the reliability of the analysis results. Medium- and complex texts often contain intricate logical relationships and specialized knowledge that a single agent cannot fully understand and process. Multiple agents can leverage their respective knowledge reserves and analytical capabilities to collectively address these challenges and improve the effectiveness of analyzing complex texts.
[0081] By rationally allocating the number of analytical agents based on text level, excessive agents are avoided when processing simple text, thus reducing computational resource consumption and costs. Simultaneously, for complex text, although the number of agents increases, the overall resource utilization efficiency is improved and the unit analysis cost is reduced because the analysis tasks are completed more efficiently. This method helps optimize resource allocation, enabling limited analytical agent resources to better meet the analysis needs of texts at different levels. Even with limited resources, resources can be rationally allocated according to the actual situation of the text, ensuring that important and complex texts receive sufficient analysis and processing.
[0082] This method offers a degree of flexibility, allowing the values of M, N, and P to be adjusted based on specific circumstances to adapt to different application scenarios and analytical needs. For instance, in scenarios requiring high precision in text analysis, the number of agents at each level can be appropriately increased. As business grows and the volume of text increases, the total number of analytical agents can be increased to expand the system's processing capacity. Furthermore, the mechanism of dynamically invoking agents based on text level ensures that the system can still utilize resources efficiently and rationally during expansion.
[0083] In this implementation, the analytical agent is preferably a large model. Furthermore, the analytical agent can also be an expert or a large model built based on expert information.
[0084] Specifically, expert profiles can be built based on each expert's professional field, rating history, available time, and workload. When real experts are not required to participate in the analysis, corresponding analytical agents can be built based on the expert profiles.
[0085] When expert participation in analysis is required, an expert database can be established first. This database records detailed information for each expert, including: their professional field (e.g., cardiovascular diseases, endocrine diseases, nervous system diseases), expert authority coefficient (range 0.1-1.0), daily workload limit (e.g., a maximum of 10 articles reviewed per day), available time slots (e.g., 9:00-12:00 AM, 2:00-6:00 PM), and historical rating quality index. Upon initiating the analysis, the expert matching algorithm is activated. This algorithm employs a multi-factor comprehensive scoring mechanism: 1) Professional matching degree calculation: The system calculates the similarity between the article's technical field and the expert's professional field using a cosine similarity algorithm. It converts domain keywords into vectors and calculates the cosine of the angle between the vectors. 1) Higher similarity, higher matching degree; 2) Workload assessment: The system monitors the current workload of each expert in real time, including the number of assigned but unfinished review tasks, the estimated completion time, etc., and prioritizes experts with lighter workloads to ensure the balance of task allocation; 3) Time availability check: The system selects experts who can complete the review within the required time based on the urgency of the article and the available time period of the experts; 4) Expert number determination rules: For low-complexity articles, 3 experts are assigned; for medium-complexity articles, 5 experts are assigned; for high-complexity articles, 7 experts are assigned. The system also considers the importance level of the article, and the number of experts can be appropriately increased for important articles.
[0086] For S300, “obtaining content quality analysis information, language expression analysis information, and structural arrangement analysis information based on each analytical agent”, the content quality analysis information includes scoring the sufficiency of arguments (0-15 points), logical rigor (0-15 points), and professional accuracy (0-10 points) respectively, and then automatically calculating the total content quality score.
[0087] The language expression analysis includes three sub-items: word accuracy (0-10 points), clarity of expression (0-10 points), and language standardization (0-10 points), and then automatically calculates the total content quality score. The structure and arrangement analysis covers logical structure (0-15 points), paragraph arrangement (0-10 points), and overall layout (0-5 points), and then automatically calculates the total content quality score.
[0088] Each scoring item provides detailed scoring criteria and examples to ensure consistency in scoring. It can also output specific review comments, explaining the reasons for deductions and suggestions for improvement.
[0089] Reference Figure 3 The method for S400, which "obtains the similarity between target information and each analytical agent, and obtains the confidence level of each analytical agent based on the similarity," includes the following methods for obtaining the confidence level of each analytical agent:
[0090] S410, acquire historical training data for each analytical agent;
[0091] S420, obtain the cosine similarity between the target information and historical training data;
[0092] S430: Based on the cosine similarity and the preset similarity scoring table, obtain the confidence level of the analysis agent corresponding to the similarity.
[0093] Specifically, a cosine similarity of [0, 0.2) corresponds to a confidence level of 0.1; a cosine similarity of [0.2, 0.4) corresponds to a confidence level of 0.3; a cosine similarity of [0.4, 0.6) corresponds to a confidence level of 0.5; a cosine similarity of [0.6, 0.8) corresponds to a confidence level of 0.7; and a cosine similarity of [0.8, 1] corresponds to a confidence level of 0.9.
[0094] In this embodiment, by acquiring the historical training data of each analytical agent and calculating the cosine similarity between the target information and the historical training data, the agent's ability to process the current target information can be evaluated from its past performance. The historical training data reflects the knowledge domain and patterns that the agent learns and adapts to. If the target information is highly similar to the historical training data, it indicates that the agent has rich experience in processing this type of information, and its analysis results are more reliable.
[0095] Based on a pre-defined similarity scoring table, cosine similarity is converted into confidence scores, which are then used to measure the reliability of each analytical agent. This allows for the subsequent use of the analytical agent's results to be filtered, weighted, or comprehensively judged based on the confidence scores, avoiding blind reliance on unreliable analytical results and thus improving the overall reliability of the analytical results.
[0096] Different target information may differ in content, structure, and domain. By calculating similarity and confidence, each analytical agent can be assigned a corresponding confidence level for different types of target information, enabling the system to flexibly select appropriate analytical agents based on the characteristics of the target information. For example, for target information in the medical field, agents with high similarity and high confidence levels to the target information in historical medical training data will be given priority.
[0097] As time goes by and the agents are continuously trained, their performance may change. By continuously calculating similarity and confidence, we can monitor the adaptability and reliability of each agent to different target information in real time, promptly detect changes in agent performance, and make corresponding adjustments to ensure the adaptability and flexibility of the system.
[0098] Furthermore, the analysis results of the content quality dimension, language expression dimension, and structural arrangement dimension can be compared with the target content score, target language score, and target structural score to obtain the differential analysis results, which include the gap in content quality, language expression, and structural arrangement.
[0099] Based on the size of the gap and the difficulty of improvement, the dimensions are prioritized, with higher priority given to dimensions that have larger gaps and are relatively easier to improve.
[0100] Furthermore, personalized optimization strategies can be generated based on the results of differential analysis. Specifically, based on the current rating level, relevant templates are retrieved from the corresponding strategy library, and the matching process considers factors such as article topic, professional field, and target audience. For articles lacking sufficient content quality, the system analyzes the existing content structure and identifies missing key elements. For example, for articles lacking case support, it is recommended to "add 2-3 typical clinical cases, each including symptom description, diagnostic process, treatment plan, and effect evaluation." The language characteristics of the article are analyzed to identify expression problems. For inappropriate terminology, it is recommended to "standardize 'stomach ache' to 'abdominal pain' and refine 'heart disease' to specific disease names such as 'coronary heart disease' or 'arrhythmia'." The logical structure of the article is evaluated to identify structural problems. For articles with disordered logic, it is recommended to "reorganize the content according to the logical order of 'disease overview - etiological analysis - symptom presentation - diagnostic methods - treatment plan - preventive measures'."
[0101] Furthermore, this application also includes establishing a dynamic feedback mechanism to continuously optimize the strategy generation effect. Specifically, when a user modifies an article based on suggestions and resubmits it for review, the changes in scores before and after the modifications are recorded, and the actual effect of each suggestion is calculated. Multi-dimensional evaluation indicators are used, including the magnitude of improvement (i.e., the degree of score improvement after modification), implementation difficulty (i.e., the proportion of users who adopt the suggestions), and time efficiency (i.e., the time from suggestion to completion of improvement). Based on the effect evaluation results, the system dynamically adjusts the weight of strategy templates; strategy templates with good effects have increased weight and are given priority in subsequent matching; strategy templates with poor effects have decreased weight.
[0102] Furthermore, we can periodically analyze the common characteristics of effective strategies to generate new strategy templates, while eliminating outdated templates that have not performed well in the long run.
[0103] Furthermore, this application also includes: supporting multiple rounds of optimization and iteration for articles; specifically, establishing a version management mechanism for each article to record the content, time, and adopted suggestions of each modification, allowing users to view the complete evolution of the article; for resubmitted articles, the system adopts an incremental review mechanism, focusing on the quality improvement of the modified parts while checking whether the modifications have introduced new problems; and setting progressive goals for each round of optimization. The first round focuses on solving basic problems, the second round focuses on refining details, and the third round pursues excellent quality.
[0104] Furthermore, this application also includes: establishing a multi-layered quality assurance mechanism; specifically, real-time monitoring of scoring consistency; automatically triggering a review mechanism when significant scoring discrepancies are detected, inviting senior experts or more specialized large-scale models to conduct arbitration scoring; employing statistical methods to detect abnormal scoring behavior, such as a single expert's score significantly deviating from other experts' scores or unusually short scoring times; and regularly providing experts with scoring quality reports, including indicators such as scoring consistency and suggestion adoption rate, to help experts improve review quality. For frequently accessed data such as expert information and strategy templates, a multi-level caching mechanism is used to reduce database access frequency; and load balancing technology is employed to rationally distribute review tasks across different servers to avoid single-point overload.
[0105] Secondly, this disclosure also provides a text content analysis system for executing the text content analysis method disclosed in the first aspect of this application. The system specifically includes:
[0106] The information extraction module is used to parse the text to be analyzed and extract target information, including keywords, technical fields, and text complexity.
[0107] The text level determination module is used to determine the text level based on the target information;
[0108] The agent invocation module is used to invoke a corresponding number of analytical agents based on the text level; and to obtain content quality analysis information, language expression analysis information, and structural arrangement analysis information based on each analytical agent.
[0109] The confidence acquisition module is used to obtain the similarity between the target information and each analysis agent, and to obtain the confidence of each analysis agent based on the similarity.
[0110] The multi-dimensional analysis module is used to obtain the content quality dimension analysis results, language expression dimension analysis results, and structure arrangement dimension analysis results based on the content quality analysis information, language expression analysis information, structural arrangement analysis information, and confidence scores of all analysis agents.
[0111] The content analysis results acquisition module is used to obtain text content analysis results based on the analysis results of content quality dimension, language expression dimension, and structure and arrangement dimension.
[0112] A computer device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.
[0113] The processor may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of this disclosure, the processor is used to execute computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the text content analysis methods of the foregoing embodiments of this disclosure.
[0114] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.
[0115] like Figure 4 This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 4 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0116] like Figure 4 As shown, a computer device may include a processor (such as a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0117] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 4 A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.
[0118] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the text content analysis method of embodiments of this disclosure are performed.
[0119] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.
[0120] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the text content analysis methods of the foregoing embodiments of the present disclosure are performed.
[0121] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).
[0122] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.
[0123] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0124] In this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, devices, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc., are open-ended terms meaning "including but not limited to," and are used interchangeably with them. The terms "or" and "and" as used herein refer to the terms "and / or," and are used interchangeably with them unless the context clearly indicates otherwise. The term "such as" as used herein refers to the phrase "such as but not limited to," and is used interchangeably with it.
[0125] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.
[0126] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.
[0127] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.
[0128] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0129] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A text content analysis method, characterized in that, include: The text to be analyzed is parsed to extract target information, which includes keywords, technical fields, and text complexity. Determining the text level based on the target information specifically includes: determining the text's professional depth score based on the number of keywords and the coverage of corresponding professional fields; determining the field's popularity and cross-disciplinary nature based on the technical field, and determining the text's difficulty score based on the field's popularity and cross-disciplinary nature; weighting and summing the text's professional depth score, text difficulty score, and text complexity according to preset weights to obtain a target comprehensive score; and determining the text level based on the target comprehensive score. The corresponding number of analytical agents are invoked based on the text level; the analytical agents are large models. Obtain content quality analysis information, language expression analysis information, and structural arrangement analysis information for each of the analytical agents; Obtaining the similarity between the target information and each of the analytical agents, and obtaining the confidence level of each analytical agent based on the similarity, specifically includes: obtaining historical training data of each analytical agent; obtaining the cosine similarity between the target information and the historical training data; and obtaining the confidence level of the analytical agent corresponding to the similarity based on the cosine similarity and a preset similarity scoring table. Based on the content quality analysis information, language expression analysis information, structural arrangement analysis information, and confidence level corresponding to all the analytical agents, the content quality dimension analysis results, language expression dimension analysis results, and structural arrangement dimension analysis results are obtained. Based on the analysis results of the content quality dimension, the language expression dimension, and the structure and arrangement dimension, the text content analysis results are obtained.
2. The text content analysis method according to claim 1, characterized in that, The method for obtaining the text complexity includes: Based on the text to be analyzed, obtain the article's word count, number of technical terms, total number of words, total number of sentences, as well as the word count, number of compound sentences, and number of cited references for each sentence; Based on the preset word count range scoring rules, the word count dimension score corresponding to the word count of the article is obtained; The terminology density is obtained based on the number of terminology and the total number of words, and the terminology density dimension score corresponding to the terminology density is obtained according to the preset terminology density interval scoring rules. Identify all sentences that exceed the preset character limit and record them as long sentences; The sentence density is obtained based on the number of complex sentences, the total number of long sentences, and the total number of sentences; According to the preset scoring rules for sentence density intervals, the sentence density dimension score corresponding to the sentence density is obtained; According to the preset scoring rules for the cited literature interval, the literature citation dimension score corresponding to the number of cited literatures is obtained; The text complexity is obtained by weighting and summing the scores for the word count dimension, the terminology density dimension, the sentence structure density dimension, and the citation dimension based on preset weights.
3. The text content analysis method according to claim 2, characterized in that, The text level can be any one of the following: normal, medium, or complex.
4. The text content analysis method according to claim 3, characterized in that, The step of calling the corresponding number of analytical agents based on the text level includes: When the text level is normal, the text is sent to M analytical agents; When the text level is medium, the text is sent to N analytical agents; When the text level is complex, the text is sent to P analytical agents; Where P > N > M ≥ 1.
5. The text content analysis method according to claim 1, characterized in that, The content quality dimension analysis result is N: ; The result of the language expression dimension analysis is Y: ; The structural arrangement dimension analysis results are as follows: : ; in, For the first Content quality analysis information corresponding to each analytical agent. For the first The confidence level of each analytical agent, where n is the total number of analytical agents. For the first The language expression analysis information corresponding to each of the aforementioned analytical agents. For the first The structural arrangement analysis information corresponding to each analytical agent.
6. A text content analysis system, characterized in that, include: The information extraction module is used to parse the content of the text to be analyzed and extract target information, including keywords, technical fields, and text complexity. The text level determination module is used to determine the text level based on the target information, specifically including: determining the text professional depth score based on the number of keywords and the coverage of the corresponding professional fields; determining the field popularity and field crossover based on the technical field, and determining the text difficulty score based on the field popularity and field crossover; weighting and summing the text professional depth score, the text difficulty score, and the text complexity according to preset weights to obtain a target comprehensive score; and determining the text level based on the target comprehensive score. The agent invocation module is used to invoke a corresponding number of analytical agents according to the text level; to obtain content quality analysis information, language expression analysis information, and structural arrangement analysis information for each analytical agent; the analytical agent is a large model; The confidence acquisition module is used to acquire the similarity between the target information and each of the analysis agents, and to acquire the confidence of each analysis agent based on the similarity. Specifically, it includes: acquiring the historical training data of each analysis agent; acquiring the cosine similarity between the target information and the historical training data; and obtaining the confidence of the analysis agent corresponding to the similarity based on the cosine similarity and a preset similarity scoring table. The multi-dimensional analysis module is used to obtain content quality dimension analysis results, language expression dimension analysis results, and structural arrangement dimension analysis results based on the content quality analysis information, language expression analysis information, structural arrangement analysis information, and confidence scores corresponding to all the analysis agents. The content analysis result acquisition module is used to obtain text content analysis results based on the content quality dimension analysis results, the language expression dimension analysis results, and the structure and arrangement dimension analysis results.
7. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the text content analysis method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the text content analysis method according to any one of claims 1-5.
9. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
An English article quality evaluation method and system
CN109670184A
Article quality evaluation method, article recommendation method and corresponding devices
CN111488931A