Method and apparatus for evaluating the quality of a presentation manuscript

The method employs a neural network to evaluate the quality of a speech manuscript by identifying key point positions and content alignment, addressing inefficiencies and subjective errors in existing methods while ensuring accurate and efficient quality assessment.

JP7693030B2Active Publication Date: 2025-06-16BEIJING UMU TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023577907
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-11-25
Filing Date
2023-04-19
Publication Date
2025-06-16
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

Existing methods for evaluating the quality of a speech manuscript are inefficient and prone to subjective errors, as they primarily focus on the logical structure without considering the content's intent and key point alignment.

Method used

A method utilizing a neural network model to semantically identify text units in a speech manuscript, determine the hit positions of key point contents, and calculate an evaluation result based on the alignment of key points and their intended order.

Benefits of technology

This approach allows for an accurate evaluation of the speech manuscript's quality by ensuring the logical structure and content alignment, reducing subjective bias and improving efficiency without requiring extensive training on specific field data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693030000001
    Figure 0007693030000001
  • Figure 0007693030000002
    Figure 0007693030000002
  • Figure 0007693030000003
    Figure 0007693030000003
Patent Text Reader

Abstract

The present application provides a method and device for evaluating the quality of a manuscript for a lecture, the method including: obtaining a manuscript for a lecture and a predetermined sequence of key points, the predetermined sequence of key points including a plurality of key points; dividing the manuscript for a lecture into a plurality of text units; performing semantic identification of the plurality of text units using a neural network model to determine a hit position in the manuscript for each of the key points; and calculating an evaluation result of the manuscript for a lecture based on the hit position and the sequence of the plurality of key points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and more specifically to a method and apparatus for evaluating the quality of a speech manuscript.

Background Art

[0002] A speech is a language communication activity in which a public place mainly uses spoken language and body language as an auxiliary means to clearly and completely express its opinions and claims on a specific issue, explain the principles, express emotions, and conduct publicity and agitation. A speech manuscript is the basis of a speech, and it must contain certain key points.

[0003] A speech manuscript always expresses the key points to be expressed in a certain order by methods such as an arrangement method or a method of generalizing first and then elaborating. Such an order may be used as an indicator for evaluating the quality of the speech manuscript. It is obvious that manually evaluating the quality of a speech manuscript is relatively inefficient and is easily affected by subjective factors of the evaluator.

[0004] Chinese Patent Document CN113361275A discloses a method for evaluating the logical structure of a speech manuscript. The method evaluates the overall logical structure by identifying the conjunctions appearing in the speech manuscript and their distribution states. Although this solution can more accurately evaluate the logic of the speech manuscript, it ignores the content that the speech manuscript intends to express, and it cannot fully show that the quality of a speech manuscript is relatively high only by reasonably using conjunctions.

Summary of the Invention

Means for Solving the Problems

[0005] In view of this, the present invention provides a method for evaluating the quality of a speech manuscript, which includes obtaining a speech manuscript and a predetermined key point sequence, where the predetermined key point sequence includes a plurality of key point contents, dividing the speech manuscript into a plurality of text units, semantically identifying the plurality of text units by a neural network model to determine the hit positions of each key point content in the speech manuscript, and calculating an evaluation result of the speech manuscript based on the hit positions and the order of the plurality of key point contents.

[0006] As an option, the step of semantically identifying the plurality of text units by a neural network model to determine the hit positions of each key point content in the speech manuscript specifically includes: the neural network model respectively identifying the text units that match each key point content, and obtaining the numbers of the text units that match each key point content.

[0007] As an option, the step of determining the evaluation result of the speech manuscript based on the hit positions specifically includes: determining the length of the valid content based on the numbers of the text units that match the first key point content and the last key point content, determining the desired positions of each key point content based on the number of key point contents and the length of the valid content, respectively determining whether the hit position of each key point content matches the desired position, and calculating the evaluation result of the speech manuscript based on all the determination results.

[0008] As an option, the length of the valid content is end - start, where start is the number of the text unit that matches the first key point content, and end is the number of the text unit that matches the last key point content.

[0009] As an option, the desired position of the key point content is a desired interval determined based on the number of key point contents, the length of the valid content, and the number of each key point content.

[0010] As an option, the desired section of the i-th key content is [start + (i - 1) * (end - start) / n, start + i * (end - start) / n], where start is the number of the text unit that matches the first key content, end is the number of the text unit that matches the last key content, and n is the number of the key content.

[0011] As an option, the method further includes obtaining an involvement command of the desired section for adjusting the desired section of any of the key contents manually input.

[0012] As an option, specifically, to determine whether the hit position of each key content matches the desired position, it includes determining whether the number of the text unit corresponding to the i-th key content is within the section, and when the number of the text unit corresponding to the i-th key content is within the section, determining that the hit position matches the desired position.

[0013] As an option, specifically, to calculate the evaluation result of the speech manuscript based on all the judgment results, in the case of a judgment result where the conclusions do not match, it includes determining a penalty amount based on the degree of deviation of the hit position from the desired position, and calculating the evaluation result of the speech manuscript based on all the penalty amounts.

[0014] The present invention further provides a quality evaluation device for a speech manuscript, including a processor and a memory connected to the processor. The memory stores instructions that can be executed by the processor. When the instructions are executed by the processor, the processor is caused to execute the quality evaluation method of the speech manuscript described above.

Effect of the Invention

[0015] According to the method and apparatus for evaluating the quality of a manuscript for a speech according to an embodiment of the present invention, first, an evaluator is allowed to provide the key points to be presented in the predicted manuscript for the speech. This solution sequentially determines the positions in the manuscript for the speech where each key point is mentioned by means of a neural network model, and evaluates whether the logical structure of the manuscript for the speech is appropriate and whether it covers all the key points to be presented according to the order of the key points and the order of the positions where they are mentioned. Based on the content that the manuscript for the speech intends to express as the main basis for judgment, the quality of the manuscript for the speech can be accurately evaluated. Moreover, this solution does not require training the neural network model with a large amount of language materials in a specific field. By using a semantic discriminable model trained with open-source data in a general field, the identification of the hit positions of the key point content can be realized, and it has higher scalability and practicality.

Brief Description of the Drawings

[0016] To more clearly explain the specific embodiments of the present invention or the technical solutions of the prior art, the following briefly describes the drawings necessary for the description of the specific embodiments or the prior art. Obviously, the drawings described below are examples of the embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative effort.

Figure 1

Figure 2

Figure 3

Modes for Carrying Out the Invention

[0017] The technical solution of the present invention will be clearly and completely described below with reference to the drawings. Of course, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, any other embodiments obtained by those skilled in the art without inventive labor belong to the protection scope of the present invention.

[0018] Embodiments of the present invention provide a method for evaluating the quality of a speech manuscript, and this method may be executed by electronic devices such as a computer and a server. As shown in FIG. 1, this method includes the following steps S1 to S4.

[0019] S1 Obtain a speech manuscript and a predetermined key point sequence, where the predetermined key point sequence includes a plurality of key point contents. These key point contents are in order, and this speech manuscript is a summary of the content expected to be expressed. For example, it may be a single word, a collocation, or a word.

[0020] An evaluator can set a plurality of key point contents based on elements such as the theme and the audience of the speech manuscript. For example, when there are n key point contents, the evaluator expects that the evaluated speech manuscript contains text content related to these n key point contents.

[0021] S2 Divide the speech manuscript into a plurality of text units. It may be divided with reference to the method described in CN113361275A, or it may be divided in a simpler way. For example, it is divided by punctuation marks indicating the end of a single sentence such as a period, a question mark, or an exclamation mark, and the text unit is a single sentence.

[0022] The S3 neural network model semantically identifies a plurality of text units to determine the hit positions in the speech manuscript of each key point content. The neural network model in this solution means should identify whether the content expressed by each word matches each predetermined key point, and specific algorithms such as a zero-resource classification model and a similarity judgment model may be used. The predetermined key point content is general and not the initial characters manually extracted from the speech manuscript. As an example, for instance, when the speech manuscript to be evaluated is related to the recommendation of electronic products, one predetermined key point may be "the hardware performance of electronic products". In this case, the neural network model should not identify whether this word exists in the speech manuscript, but should identify each text unit and determine whether its meaning matches the hardware performance of electronic products.

[0023] Assuming there are m text units and n key point contents, the neural network model sequentially determines the positions where each key point is hit in the speech manuscript. As shown in Figure 2, for example, for the i-th key point content, the neural network model determines whether the content expressed by text unit m among the m text units i … text unit m j matches the i-th key point content. If so, the position of text unit m i … text unit m j in the entire speech manuscript is the hit position of the i-th key point content in the speech manuscript.

[0024] Note that the number of text units explaining a certain key point content in a speech manuscript may be one or more, and there may also be a possibility that the text units explaining the key point content are not identified, that is, there is no text expressing the key point content in the entire speech manuscript.

[0025] Calculate the evaluation result of the speech manuscript based on the hit position in S4 and the order of multiple key points. The core idea of this solution is that a high-quality speech manuscript should meet the requirement that the order of a predetermined key point and the order of the hit key points in the speech manuscript have a linear correlation. For example, the hit position of the i-th key point content should be before the hit position of the (i + 1)-th key point content and after the hit position of the (i - 1)-th key point content. If the hit positions of all key point contents conform to the above relationship, a better evaluation result can be obtained. On the contrary, if the above relationship is not met, or some key point contents do not have hit positions, a worse evaluation result will be obtained.

[0026] The evaluation result may be a numerical value or a classification result such as excellent, good, pass, and poor. There are various specific implementation logics. The more cases where it does not conform to the above linear relationship, the larger the gap and the more negative situations such as not being hit. Then, the score will be lower or the classification result will be worse. In the opposite case, the score will be higher or the classification result will be better. Therefore, penalties can be calculated for negative situations or incentives can be calculated for positive situations, and the evaluation result can be obtained based on the penalties or incentives.

[0027] In a preferred embodiment, in step S3, the neural network model respectively identifies the text units that match each key point content, and further obtains the numbers of the text units that match each key point content. Assuming that a speech manuscript is divided into m words, the numbers of m text units are obtained, and the number of the text unit that matches the i-th key point content is the hit position of the i-th key point content.

[0028] Furthermore, in step S4, based on the numbers of the text units that match the first key point content and the last key point content, the length of the effective content is determined. Since the beginning and the end of the manuscript for the speech are generally characters that have nothing to do with all the key point contents, in order to accurately evaluate whether the distribution of the hit positions of each key point content is uniform, the length of the effective content is first determined here. Denote the number of the text unit that matches the first key point content as start, and denote the number of the text unit that matches the last key point content as end. Then, the length of the effective content is end - start.

[0029] Based on the number of the key point contents and the length of the effective content, the desired positions of each key point content are determined. In order to determine whether the hit positions are evenly distributed, it is necessary to determine the desired positions in combination with the total number of the predetermined key point contents. For example, when the number of the predetermined key points is relatively small, for example, there are only 3, it is reasonable that the hit position of the first predetermined key point is in the upper one-third of the effective content of the manuscript for the speech, and the upper one-third of the effective content is the desired position of the first predetermined key point. Similarly, the desired position of the second predetermined key point is in the middle one-third of the effective content, and the desired position of the third predetermined key point is in the last one-third of the effective content. When the number of the predetermined key points is relatively large, for example, there are 10, the desired positions of the predetermined key points are also adjusted accordingly.

[0030] Respectively determine whether the hit positions of each key point content match the desired positions, and calculate the evaluation result of the manuscript for the speech based on all the judgment results. The situation where the hit position is not at its expected position is a negative situation. The farther it deviates from the expected position, the worse the judgment result for the key point content, and the penalty value may be set to increase accordingly. By summarizing the judgment results of the hit positions of all the key point contents, the comprehensive evaluation result can be calculated.

[0031] Furthermore, the desired position of the above key point content is a desired interval determined based on the number of key point contents, the length of the valid content, and the numbers of each key point content. The length of the desired interval of each predetermined key point may be shown as (end - start) / n, and the desired interval of the i-th key point content is [start + (i - 1) * (end - start) / n, start + i * (end - start) / n], where start is the number of the text unit that matches the first key point content, end is the number of the text unit that matches the last key point content, and n is the number of key point contents.

[0032] The device for executing this method by the above preferred solution can automatically determine the desired interval of each key point content according to the manuscript for the speech and the actual situation of the predetermined key point content. In an alternative embodiment, it may also be allowed to manually adjust the above desired interval.

[0033] For example, if some key point contents are more important or less important than other key point contents, the content expressing the more important key point content in the manuscript for the speech should be longer, and thus its desired interval should be longer. Conversely, in the opposite case, the desired interval should be shorter. And the length of the automatically determined desired interval is an average. In a preferred embodiment, in order to obtain a more accurate evaluation result, first, the desired intervals of all key point contents are automatically obtained, and then a participation command for the desired interval for manually inputting and adjusting the desired interval of any key point content to diverge to both ends may be obtained. It is possible to shorten or extend any desired interval as needed.

[0034] After obtaining the desired interval of each key point content, it is determined whether the number of the text unit corresponding to the i-th key point content is within the interval. If the number of the text unit corresponding to the i-th key point content is within the interval, it is determined that the hit position coincides with the desired position. For example, if the number of the text unit corresponding to the i-th key point content is m i …m j , then m i …mj ∈[start+(i - 1) * (end - start) / n, start + i * Determine whether [(end - start) / n, start + i] contains [start+(i - 1), (end - start) / n].

[0035] In the case of a judgment result where the conclusions do not match, that is, when the above conditions are not met, determine the penalty amount based on the degree of deviation of the hit position from the desired position. The farther the hit position deviates from the expected interval, the larger the penalty amount. In a specific example, the penalty amount is defined as the percentage of the length of deviation from the desired interval in the total length of the effective content.

[0036] Finally, calculate the evaluation result of the speech manuscript based on all the penalty amounts. Figure 3 shows a specific example. The vertical axis shows the predetermined key points 1 to n from top to bottom, and the horizontal axis shows the text units 1 to m from left to right. The colored patches in it show the hit situation of the key points in the text unit. The dark - colored patches indicate that the text unit matches the key point, and the light - colored patches indicate non - matching. The identification result by the neural network model is the degree of coincidence. In the legend, the darker the color, the higher the degree of coincidence of the colored patch, and vice versa. The content in the dotted - line area indicates that the order of matching the predetermined key points and the order of hitting the key points in the speech manuscript have a linear correlation. Outside the dotted - line area, it shows the position of the text unit hitting the key point in the speech manuscript. If this position does not match the order of the key point itself, a penalty amount will be generated, and finally, the final result can be calculated based on all the penalty amounts.

[0037] As will be understood by those skilled in the art, embodiments of the present invention may be provided as a method, system, or computer program product. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Further, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc.) having computer-usable program code embodied therein.

[0038] The present invention will be described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions may be provided to the processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus for manufacturing a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for realizing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0039] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory create an article of manufacture including instruction means for realizing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0040] These computer program instructions, when loaded into a computer or other programmable data processing apparatus, may cause the computer or other programmable device to execute a series of operational steps so as to generate computer-executable processing, whereby the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one process or a plurality of processes in a flowchart and / or one block or a plurality of blocks in a block diagram.

[0041] Obviously, the above embodiments are merely examples given for clear description and do not limit the embodiments. Those skilled in the art can make further different forms of changes or variations based on the above description. Here, all embodiments cannot be listed, nor is it necessary. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.

Claims

1. A method for evaluating the quality of a speech manuscript, comprising: Step S1 in which a processor obtains a speech manuscript and a predetermined key point sequence, where the predetermined key point sequence includes a plurality of key point contents; Step S2 in which the processor divides the speech manuscript into a plurality of text units; Step S3 in which the processor semantically identifies the plurality of text units using a neural network model and determines the hit positions of each key point content in the speech manuscript; Step S4 in which the processor calculates an evaluation result of the speech manuscript based on the hit positions and the order of the plurality of key point contents, characterized in that it includes the above steps.

2. Step S3 of semantically identifying the plurality of text units using a neural network model and determining the hit positions of each key point content in the speech manuscript is specifically: The processor respectively identifies the text units that the neural network model matches with each key point content; The processor obtains the numbers of the text units that match with each key point content, characterized in that it includes the above steps according to Claim 1.

3. Step S4 of determining the evaluation result of the speech manuscript based on the hit positions is specifically: The processor determines the length of the valid content based on the numbers of the text units that match with the first key point content and the last key point content; The processor determines the desired positions of each key point content based on the number of key point contents and the length of the valid content; The processor respectively determines whether the hit positions of each key point content match the desired positions; The method according to claim 2, further comprising: the processor calculating an evaluation result of the speech manuscript based on all the determination results.

4. The length of the effective content is end - start, where start is the number of the text unit that matches the first key point content, and end is the number of the text unit that matches the last key point content. The method according to claim 3, characterized in that.

5. The desired position of the key point content is a desired interval determined based on the number of the key point content, the length of the effective content, and the number of each key point content. The method according to claim 3, characterized in that.

6. The desired interval of the i - th key point content is [start+(i - 1) * (end - start) / n, start + i * (end - start) / n], where start is the number of the text unit that matches the first key point content, end is the number of the text unit that matches the last key point content, and n is the number of the key point content. The method according to claim 5, characterized in that.

7. The method according to claim 5, further comprising: obtaining an intervention command of the desired interval for adjusting the desired interval of any of the key point contents manually input.

8. Determining whether the hit position of each key point content coincides with the desired position specifically includes: the processor determining whether the number of the text unit corresponding to the i - th key point content is within the desired interval of the i - th key point content; when the processor determines that the number of the text unit corresponding to the i - th key point content is within the desired interval of the i - th key point content, determining that the hit position coincides with the desired position. The method according to claim 6, characterized in that.

9. Calculating the evaluation result of the speech manuscript based on all the judgment results specifically includes: When the judgment results do not match in conclusion, the processor determines a penalty amount based on the degree of deviation of the hit position from the desired position; The processor calculates the evaluation result of the speech manuscript based on all the penalty amounts. The method according to claim 3 is characterized by including the above. **Claim 10** A quality evaluation device for a speech manuscript, comprising: A processor and a memory connected to the processor, wherein the memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor is caused to execute the quality evaluation method for a speech manuscript according to any one of claims 1 to 9. A quality evaluation device for a speech manuscript is characterized by this.

Citation Information

Patent Citations

  • Online course video resource content identification and evaluation method and intelligent system

    CN111898441A

  • Document detection method and device, equipment and storage medium

    CN113515628A

  • Presentation data retrieving system and its method and program

    JP2004265097A

  • Device, System, and Method for Assigning Differential Weight to Meeting Participants and for Generating Meeting Summaries

    US20200243095A1