Paperless real-time evaluation processing method for report and related product

Through the intelligent reporting and evaluation system, the problem that reporting personnel in traditional meetings cannot understand feedback in real time is solved, and the reporting efficiency and accuracy are improved.

WO2025118141A1PCT designated stage expired Publication Date: 2025-06-12SHENZHEN TAIDEN INDAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/136455
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In traditional paperless conference systems, reporting personnel cannot understand the feedback of participants in real time, resulting in low reporting efficiency and accuracy.

Method used

Through the intelligent reporting and evaluation system, the electronic speeches and audio of the reporter are analyzed in real time, and the interactive annotations of the participants and the feedback information of the reporter are obtained, the report evaluation is generated, and the report evaluation is sent to the reporter.

Benefits of technology

It improves the reporting efficiency and accuracy of the reporter, allowing the reporter to adjust the reporting rhythm and content after obtaining the evaluation in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023136455_12062025_PF_FP_ABST
    Figure CN2023136455_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a paperless real-time evaluation processing method for a report and a related product. The method is applied to a paperless real-time evaluation processing device for a report, the paperless real-time evaluation processing device for a report is located in a paperless real-time evaluation processing system for a report, and the paperless real-time evaluation processing system for a report comprises a first terminal and a plurality of second terminals. The method comprises: acquiring, from the first terminal, an electronic speech draft of a reporter and a first report audio; acquiring, from at least one second terminal among the plurality of second terminals, at least one interactive annotation of a participant for a report of the reporter; acquiring at least one piece of feedback information of the reporter for the at least one interactive annotation; on the basis of the electronic speech draft of the reporter, the first report audio, the at least one interactive annotation, and the at least one piece of feedback information, acquiring a report evaluation for the reporter; and sending the report evaluation to the first terminal. The embodiments of the present application can improve the reporting efficiency and accuracy of a conference.
Need to check novelty before this filing date? Find Prior Art

Description

A paperless real-time evaluation report processing method and related products Technical Field

[0001] The present application relates to the field of intelligent conference communication technology, and in particular to a paperless real-time evaluation report processing method and related products. Background Art

[0002] With the continuous development and advancement of communication technology, people's work and life are becoming more and more intelligent. Among them, the widespread application of paperless conference systems has brought convenience to enterprises' smart office.

[0003] A paperless conference system uses electronic devices and software to share, communicate, and manage meeting materials. This system enables meeting recording, document sharing, and video conferencing, improving meeting efficiency and participation while reducing costs and environmental impact.

[0004] In traditional paperless conference systems, reporters can only judge the progress of the meeting on their own, and the efficiency and accuracy of meeting reports are low.

[0005] Summary of the Invention

[0006] The embodiment of the present application provides a paperless real-time evaluation reporting processing method and related products. By analyzing the meeting minutes through an intelligent reporting evaluation system and providing real-time feedback to the reporter, the problem of the reporter being unable to understand the feedback of the participants and the overall quality of the report is solved, thereby improving the reporter's reporting efficiency and accuracy.

[0007] In a first aspect, an embodiment of the present application provides a paperless real-time evaluation report processing method, the method being applied to a paperless real-time evaluation report processing device, the paperless real-time evaluation report processing device being located in a paperless real-time evaluation report processing system, the paperless real-time evaluation report processing system further comprising a first terminal and a plurality of second terminals, wherein the first terminal is a terminal of a reporter, the plurality of second terminals are terminals of a plurality of participants, and the plurality of second terminals correspond one-to-one to the plurality of participants; the method comprising:

[0008] Obtaining from the first terminal the electronic speech manuscript of the reporter and a first report audio of the electronic speech manuscript, wherein the first report audio is the audio of the reporter during the entire report process;

[0009] obtaining, from at least one second terminal among the plurality of second terminals, at least one interactive annotation of a participant on the report of the reporter;

[0010] Obtaining at least one piece of feedback information from the reporter regarding the at least one interactive annotation;

[0011] Obtaining a report evaluation for the reporter based on the reporter's electronic speech manuscript, the first report audio, the at least one interactive annotation, and the at least one feedback information;

[0012] The report evaluation is sent to the first terminal.

[0013] In a second aspect, an embodiment of the present application provides a paperless real-time evaluation report processing device, the paperless real-time evaluation report processing device comprising an acquisition unit and a processing unit;

[0014] The acquiring unit is configured to acquire, from the first terminal, the electronic speech manuscript of the reporter and a first report audio for the electronic speech manuscript, wherein the first report audio is the audio of the reporter during the entire report process;

[0015] The processing unit is configured to obtain, from at least one second terminal among the plurality of second terminals, at least one interactive annotation of a participant on the report of the reporter;

[0016] Obtaining at least one piece of feedback information from the reporter regarding the at least one interactive annotation;

[0017] Obtaining a report evaluation for the reporter based on the reporter's electronic speech manuscript, the first report audio, the at least one interactive annotation, and the at least one feedback information;

[0018] The report evaluation is sent to the first reporting terminal.

[0019] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory, wherein the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device performs the method described in the first aspect.

[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor as described in the first aspect.

[0021] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer is operable to enable the computer to execute the method described in the first aspect.

[0022] The implementation of the embodiments of the present application has the following beneficial effects:

[0023] By obtaining the presenter's electronic speech manuscript and a first audio recording of the electronic speech manuscript, the system analyzes the presenter's report content based on the electronic speech manuscript and the first audio recording, ensuring that the quality of the presenter's report can be intelligently analyzed by the evaluation system. Furthermore, during the presentation, the system obtains at least one interactive comment from a participant regarding the presenter's report, as well as the presenter's feedback. Using all of this information, the system intelligently analyzes the report and sends the evaluation to the first terminal. This allows the presenter to obtain the evaluation system's intelligent analysis in real time, adjust the pace and content of their presentation, and improve the efficiency and accuracy of their presentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] FIG1 is a schematic diagram of a paperless real-time evaluation report processing system provided in an embodiment of the present application;

[0025] FIG2 is a flow chart of a paperless real-time evaluation report processing method provided in an embodiment of the present application;

[0026] FIG3 is a flow chart of a speech script synchronization method provided by an embodiment of the present application;

[0027] FIG4 is a flow chart of a method for generating a speech synchronization signal according to an embodiment of the present application;

[0028] FIG5 is a flow chart of a method for synchronizing an electronic speech script according to an embodiment of the present application;

[0029] FIG6 is a concurrent flow chart of a speech script synchronization event provided by an embodiment of the present application;

[0030] FIG7 is a schematic diagram of a feedback operation interface provided in an embodiment of the present application;

[0031] FIG8 is a flow chart of a reporting and evaluation method provided in an embodiment of the present application;

[0032] FIG9 is a block diagram of functional units of a paperless real-time evaluation report processing device provided in an embodiment of the present application;

[0033] FIG10 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0035] The terms "first," "second," "third," and "fourth," etc., in the specification and claims of this application and the accompanying drawings are used to distinguish different objects, not to describe a particular order. In addition, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0036] References herein to "embodiments" mean that a particular feature, result, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0037] The terminal involved in the embodiment of the present application may, for example, be a conference terminal. The conference terminal is a device for conducting offline meetings, remote meetings and video calls, and the conference terminal includes but is not limited to a camera, a microphone, a speaker and a screen. The conference terminal can be connected to a local area network (LAN), such as an enterprise intranet, or can be connected to the Internet, and realize data transmission between different conference terminals through optical fiber, satellite and media with data transmission function, and conduct video calls and file sharing with other conference terminals. The conference terminal also has video transmission and audio transmission functions, and can realize multi-party video calls and screen sharing. The conference terminal also includes a user equipment (UE) with a data interaction function with a server, and the UE may include, for example, a computer, a smart phone, a portable intelligent communication device, etc. In an embodiment of the present application, the conference terminal may be, for example, an integrated conference terminal, which is a device that supports multiple communication protocols and standards and is used to support remote meetings and offline meetings. The integrated conference terminal can be, for example, an intelligent communication device that integrates components such as a high-definition camera, a microphone, a speaker, a touch screen, a stylus and an operating system to realize functions such as video conferencing, audio conferencing, screen sharing, and file sharing. Refer to Figure 1, which is a schematic diagram of a paperless real-time evaluation report processing system provided in an embodiment of the present application. The paperless real-time evaluation report processing system includes a paperless real-time evaluation report processing device 110, a first terminal 100 and multiple second terminals, such as the second terminal 101, the second terminal 102, ..., the second terminal 10n shown in Figure 1. Among them, the first terminal 100 is the reporter's terminal, the second terminal 101, the second terminal 102, ..., the second terminal 10n is the terminal of each participant, and each second terminal corresponds to each participant one-to-one.

[0038] The paperless real-time evaluation and reporting processing device 110 initiates a meeting and creates and stores meeting information. After the first terminal 100 and each second terminal establish a connection with the paperless real-time evaluation and reporting processing device 110, the paperless real-time evaluation and reporting processing device 110 adds the first terminal 100 and each second terminal to the meeting and sends meeting information to the first terminal 100 and each second terminal. The first terminal 100 and each second terminal receive the meeting information and display the meeting interface on their respective display interfaces. Among them, the first terminal 100 is directly connected to the paperless real-time evaluation and reporting processing device 110 to realize information interaction. Each second terminal is directly connected to the paperless real-time evaluation and reporting processing device 110 to realize information interaction. The first terminal 100 and each second terminal realize information interaction between the first terminal 100 and each second terminal through information interaction with the paperless real-time evaluation and reporting processing device 110.

[0039] In an embodiment of the present application, the paperless real-time evaluation report processing device 110 can be a server. The server is a computer system used to store, process, and provide data, services, or resources, and is typically used to support network services such as website hosting, email, databases, file sharing, etc. The server can be an independent physical server (PS). A physical server refers to an actual hardware device, typically a computer specifically designed to host server software and applications. The server can also be a virtualized entity. Through virtualization technology, a physical server can be divided into multiple virtual private servers (VPS), each of which can run an independent operating system and application to host websites, applications, file storage, etc. The server can also be a cloud server (Elastic Compute Service, ECS), which is based on cloud computing technology and allocates resources to users through a virtualization platform provided by a cloud service provider. The server virtualized on the cloud platform is a cloud server, and users can expand or reduce resources at any time according to their needs. The server can also be an embedded server (Embedded Service Processor, ESP), which is typically a small server embedded in other devices or systems to provide specific services or functions, such as routers, smart home devices, etc. The server can also be a server cluster (SC), which is a cluster composed of multiple servers used to jointly process a large number of requests or provide high availability and fault tolerance.

[0040] Based on the paperless real-time evaluation and reporting processing system shown in FIG1 , the paperless real-time evaluation and reporting processing device 110 obtains the reporter's electronic speech manuscript and a first report audio for the electronic speech manuscript from the first terminal 100 , wherein the first report audio is the reporter's audio during the entire reporting process;

[0041] acquiring, from at least one second terminal among the second terminal 101 , the second terminal 102 , . . . , and the second terminal 10n , at least one interactive annotation of a participant on the report of the reporter;

[0042] Obtain at least one piece of feedback information from the reporter regarding the at least one interactive comment;

[0043] Obtaining a report evaluation of the reporter based on the reporter's electronic speech manuscript, the first report audio, the at least one interactive annotation, and the at least one feedback information;

[0044] The report evaluation is sent to the first terminal 100 .

[0045] Refer to Figure 2, which is a flow chart of a paperless real-time evaluation report processing method provided in an embodiment of the present application. The method is applied to the above-mentioned paperless real-time evaluation report processing device, which is located in the above-mentioned paperless real-time evaluation report processing system.

[0046] It should be noted that the paperless real-time evaluation report processing device can be a server. In the embodiment of the present application, the paperless real-time evaluation report processing device is taken as an example to illustrate the paperless real-time evaluation report processing method.

[0047] The method includes but is not limited to the following steps:

[0048] 201: The server obtains the reporter's electronic speech manuscript and a first report audio for the electronic speech manuscript from the first terminal.

[0049] In an embodiment of the present application, by connecting a first terminal and multiple second terminals to a server, for example, the first terminal and multiple second terminals can be connected to the server respectively via the websocket protocol. After the first terminal and multiple second terminals are connected to the server respectively, the server sends an acquisition instruction to the first terminal, the first terminal receives the acquisition instruction, and sends the electronic speech manuscript of the reporter to the server. Furthermore, the first terminal records the reporter's report audio throughout the entire report process in real time, and sends the reporter's first report audio based on the electronic speech manuscript to the server.

[0050] It should be noted that after the first terminal receives the acquisition instruction sent by the server, sending the reporter's electronic speech manuscript to the server and sending the first report audio for the electronic speech manuscript to the server is a concurrent processing process, that is, multiple cores of multiple processors send the reporter's electronic speech manuscript to the server and send the first report audio for the electronic speech manuscript to the server at the same time.

[0051] In a feasible embodiment, electronic speech manuscript synchronization can be achieved by using the flowchart of the speech manuscript synchronization method shown in Figure 3. As shown in Figure 3, the electronic speech manuscript synchronization method also includes the following steps:

[0052] 301: The first terminal sends a voice signal of the reporter to the server;

[0053] 302: The server processes the voice signal through a voice transcription module and generates a speech script synchronization signal;

[0054] 303: The server sends a speech script synchronization signal to the second terminal;

[0055] 304: The second terminal processes the speech manuscript synchronization signal through the terminal speech manuscript module to achieve speech manuscript synchronization.

[0056] Specifically, the first terminal sends the presenter's voice signal to the server. Upon receiving the signal, the server's voice transcription module transcribes it and processes it to generate a speech synchronization signal. This signal includes the presenter's current presentation location and speech playback speed. The server then sends this signal to each second terminal, which then processes it through its own speech module, synchronizing the electronic speech between the first and second terminals.

[0057] The server processes the voice signal through the voice transcription module and generates a speech synchronization signal, as shown in FIG4 , which also includes the following steps:

[0058] 401: Perform voice transcription service on the reporter's voice signal and output text;

[0059] 402: Decompose the text into words to obtain multiple first keywords;

[0060] 403: Analyze the ambiguous sound in the reporter's voice signal to obtain multiple second keywords;

[0061] 404: Searching for the plurality of first keywords and the plurality of second keywords in the speech manuscript, thereby locating the current position of the speech manuscript;

[0062] 405: Based on the reporter's voice signal and the current position of the speech, generate a playback position and playback speed at that moment;

[0063] 406: Generate a synchronization event based on the playback position and playback speed at that moment.

[0064] Specifically, after receiving the presenter's voice signal, the server outputs the text corresponding to the voice signal through the voice transcription module's transcription service. The text is then decomposed into multiple first keywords. By performing audio analysis on the ambiguous sounds in the presenter's voice signal, multiple second keywords of the presenter's voice signal are obtained. The first and second keywords are then combined as the complete set of keywords for the presenter's voice signal.

[0065] Furthermore, the electronic speech manuscript is searched for the plurality of first keywords and the plurality of second keywords. If the plurality of first keywords and the plurality of second keywords are found, the electronic speech manuscript is located at the current position of the speech manuscript; otherwise, the electronic speech manuscript is searched for the plurality of first keywords and the plurality of second keywords until the plurality of first keywords and the plurality of second keywords are found.

[0066] Furthermore, based on the current position of the speech and the presenter's voice signal, the server analyzes the presenter's speech speed and generates the playback position and playback speed of the electronic speech. Then, based on the transcription time, a synchronization event is generated to synchronize the electronic speech between the first terminal and each second terminal.

[0067] For example, referring to FIG5 , FIG5 is a flow chart of a method for synchronizing an electronic speech manuscript provided by an embodiment of the present application. The method for synchronizing an electronic speech manuscript further includes the following steps:

[0068] 501: The server obtains a synchronization request from the first terminal.

[0069] In the embodiment of the present application, the synchronization request is used to request the provision of an electronic speech script synchronization service for the reporter. When the reporter starts to report, the first terminal sends a synchronization request to the server after recognizing the reporter's voice signal, and the electronic speech script synchronization service is provided to the first terminal through the synchronization request.

[0070] 502: In response to the synchronization request, the server obtains the second reporting audio of the reporter at the current moment from the first terminal.

[0071] In an embodiment of the present application, when the server receives the synchronization request from the first terminal, it sends a response signal to the first terminal. After receiving the response signal, the first terminal sends the second report audio of the reporter at the current moment to the server, and the server performs voice recognition analysis on the second report audio.

[0072] 503: The server performs semantic analysis on the second report audio to obtain the first report position of the reporter at the current moment.

[0073] Specifically, after obtaining the second report audio of the current reporter, the server performs voice recognition on the second report audio and transcribes the second report audio into text using a voice transcription function. The transcribed text of the second report audio is then subjected to semantic analysis, and the transcribed text of the second report audio is decomposed into words based on the meaning of the words in the text to obtain multiple keywords.

[0074] Furthermore, the electronic speech manuscript is searched for the multiple keywords. If the multiple keywords are found in succession, the current position of the electronic speech manuscript at that moment is recorded, i.e., the first reporting position of the presenter at that moment. If the multiple keywords are not found in the electronic speech manuscript, the search is continued until the multiple keywords are found in the electronic speech manuscript.

[0075] Therefore, through voice recognition and voice transcription of the second report audio of the current reporter, the first report position of the current reporter can be located in the electronic speech manuscript based on the transcribed text.

[0076] 504: The server performs speech speed analysis on the second report audio to obtain the current speaking speed of the reporter.

[0077] In an embodiment of the present application, after the server performs voice recognition on the second report audio, it can perform speech speed analysis on the second report audio by transcribing the audio content reported from the previous moment to the current moment into text.

[0078] Specifically, the server extracts the report audio from the previous moment to the current moment from the second report audio. For example, the report audio is the audio 0.5 seconds before the current moment in the second report audio. By performing phonetic transcription on the report audio, multiple characters based on the report audio are obtained. Furthermore, by performing voice recognition and matching on the ambiguous sounds in the report audio, all characters corresponding to the report audio are obtained.

[0079] Specifically, the multiple characters obtained by transcribing the report speech are matched with the electronic speech manuscript. If the multiple characters are found in the electronic speech manuscript, the current position is located, and the ambiguous sound in the report speech is identified and matched with the character before or after the current position. If the match is successful, the corresponding characters are added to the multiple characters to obtain all the characters corresponding to the report speech; otherwise, the multiple characters are searched again in the electronic speech manuscript. If the multiple characters are not found in the electronic speech manuscript, the search continues until the multiple characters are found in the electronic speech manuscript.

[0080] Furthermore, the speaking speed of the reporter at the current moment, that is, the number of characters reported per unit time, can be obtained by the ratio of all characters corresponding to the reported voice obtained above to the time length from the previous moment to the current moment.

[0081] 505: The server predicts the second reporting position of the reporter at the next moment based on the reporter's speaking speed at the current moment, the reporter's first reporting position, and the electronic speech manuscript.

[0082] In the embodiment of the present application, after obtaining the first reporting position of the reporter at the current moment and the speaking speed of the reporter at the current moment, the server matches them with the content of the electronic speech manuscript.

[0083] Specifically, the server locates the first reporting position of the electronic speech at the current moment. Based on the current reporter's speaking speed (i.e., the number of characters reported per unit time), the server moves backward from the first reporting position of the electronic speech by [(reporter's speaking speed) * (next moment - current moment)] characters to locate the position where the reporter is likely to report at the next moment, i.e., the predicted second reporting position of the reporter at the next moment.

[0084] 506: The server determines the number of words that the reporter needs to report from the current moment to the next moment based on the second reporting position and the first reporting position.

[0085] In an embodiment of the present application, after predicting the second reporting position of the reporter at the next moment, the server performs semantic analysis on the character text content from the first reporting position to the second reporting position in the electronic speech manuscript, and breaks down the character text into multiple words based on semantics, where the number of characters in each word can be determined according to different semantics.

[0086] It should be understood that since the report content may contain words that do not have complete meaning but have grammatical meaning or function, such as conjunctions, prepositions, modal particles, etc., when performing semantic analysis on the above character text, the number of words that do not have complete meaning but have grammatical meaning or function is recorded.

[0087] The server then counts the number of content words in the aforementioned words and the number of words that lack complete meaning but have grammatical meaning or function. The number of words that lack complete meaning but have grammatical meaning or function is multiplied by a weight, for example, 0.75, and the sum is added to the number of content words in the aforementioned words to obtain the number of words that the reporter needs to report from the current moment to the next moment.

[0088] 507: The server determines the playback speed of the electronic speech based on the number of words and the time between the current moment and the next moment.

[0089] In an embodiment of the present application, after the server obtains the number of words that the reporter needs to report from the current moment to the next moment, it can obtain the number of words that need to be played in the electronic speech script per unit time by the ratio of the number of words to the time length between the current moment and the next moment.

[0090] Furthermore, the server obtains the page display area size of the first terminal and the adaptation ratio of the display area. According to the page display area size and the adaptation ratio of the display area, it can be obtained how many lines the page display area can display and how many words each line can display.

[0091] Furthermore, the ratio of the number of words that need to be played in the electronic speech per unit time to the words that can be displayed in each line of the page display area can be used to determine the number of lines that can be displayed in the page display area per unit time of the first terminal, that is, the playback speed of the electronic speech.

[0092] 508: The server synchronizes the playback speed and the first reporting position to the first terminal and the plurality of second terminals, so that the first terminal and the plurality of second terminals all play the electronic speech manuscript from the first reporting position and at the playback speed at the current moment.

[0093] In the embodiment of the present application, after obtaining the playback speed and the first reporting position of the electronic speech manuscript, the server sends the playback speed and the first reporting position of the electronic speech manuscript to the first terminal and each second terminal that has subscribed to the speech manuscript synchronization service.

[0094] Furthermore, after receiving the playback speed and first reporting position of the electronic speech, the first terminal and each second terminal that has subscribed to the speech synchronization service will play the electronic speech at the above playback speed in the display area of ​​their respective terminals starting from the first reporting position at the current moment.

[0095] It should be noted that the page display area size and the adaptation ratio of the display area may be different between the first terminal and each second terminal, and the above-mentioned playback speed is determined based on the page display area size and the adaptation ratio of the display area of ​​the first terminal. Therefore, when each second terminal plays the electronic speech, the page display area size and the adaptation ratio of the display area of ​​each terminal are proportionally changed according to the ratio of the page display area size and the adaptation ratio of the display area of ​​each terminal to the first terminal, so as to achieve synchronization of the electronic speech between the first terminal and each second terminal that has subscribed to the speech synchronization service.

[0096] In a feasible embodiment, the first terminal and each second terminal that subscribes to the speech script synchronization service can establish a connection with the server through a websocket service. Referring to Figure 6, Figure 6 is a concurrent flow chart of a speech script synchronization event provided by an embodiment of the present application, including the following steps:

[0097] 601: The second terminal subscribes to the speech script synchronization service from the server;

[0098] 602: The server generates a synchronization event at that moment and sends the synchronization event at that moment to the data cache;

[0099] 603: The second terminal obtains the synchronization event at that moment from the data cache through the websocket service;

[0100] 604: The second terminal updates the progress of the speech based on the synchronization event at that moment.

[0101] Specifically, after the connection is established, the second terminal subscribes to the speech synchronization service from the server via the WebSocket protocol. Upon receiving the subscription request from the second terminal, the server generates a time synchronization event and obtains the first reported position and playback speed at the current moment from the first terminal. The method for obtaining the first reported position and playback speed at the current moment is the same as the steps described above and will not be described here. The server then sends the first reported position and playback speed to the data cache.

[0102] Furthermore, the server obtains the first reporting position and playback speed from the data cache via the websocket service, and transmits the first reporting position and playback speed to the first terminal and each second terminal that has subscribed to the speech script synchronization service. Thus, the first terminal and each second terminal that has subscribed to the speech script synchronization service can play the electronic speech script from the current first reporting position at the playback speed.

[0103] It can be seen that by obtaining the second report audio of the first terminal and performing semantic analysis and speech speed analysis on the second report audio, the first reporting position and speech speed of the reporter at the current moment can be analyzed, thereby predicting the second reporting position of the reporter at the next moment, and obtaining the number of words that the reporter needs to report, and then obtaining the playback speed. The playback speed and the first reporting position are synchronized to the first terminal and multiple second terminals, so that the playback progress of the electronic speech manuscript and the reporter's report content can be synchronized, and the electronic speech manuscripts of the first terminal and multiple second terminals are synchronously displayed, ensuring that the reporter can independently determine the progress of the meeting when reporting in a paperless meeting, thereby improving the reporting efficiency of the meeting.

[0104] It should be noted that during the presentation, participants' second terminals can send interactive comments to the first terminal via the server, and the presenter will provide feedback on these comments. The server's evaluation system will monitor the feedback processing, consolidate the feedback, and generate a meeting briefing.

[0105] 202: The server obtains at least one interactive annotation of a participant on the reporter's report from at least one second terminal among the plurality of second terminals.

[0106] In the embodiment of the present application, while the meeting is in progress, the display interface of each second terminal displays the feedback operation interface shown in FIG7 .

[0107] Specifically, the electronic speech manuscript is played synchronously in the display area above the display interface. The information prompt box below contains a prompt message of "input feedback information", which instructs each participant to input the information that needs to be fed back in the information prompt box regarding the content of the reporter's report. There is a send function control near the information prompt box, and each participant can send the above-input feedback information to the server by clicking the send function control. The display interface also includes an annotation function control. By clicking the annotation function control, each participant can annotate the electronic speech manuscript in the upper display area regarding the content of the reporter's report, and then click the send function control to send the annotated content to the server. The feedback information and annotation content input by the above-mentioned participants are at least one interactive annotation of the participant on the reporter's report.

[0108] Furthermore, after obtaining at least one interactive annotation of a participant regarding the presenter's report from at least one of the plurality of second terminals, the server transmits the interactive annotation to the first terminal and each of the second terminals. After receiving the at least one interactive annotation, the first terminal and each of the second terminals simultaneously display the at least one interactive annotation in the form of a bullet screen on their display interfaces.

[0109] 203: The server obtains at least one piece of feedback information from the reporter regarding the at least one interactive annotation from the first terminal.

[0110] In an embodiment of the present application, after obtaining at least one of the above-mentioned interactive annotations, the server identifies and analyzes the content of the interactive annotation, and classifies the interactive annotation based on the results of the identification and analysis. For example, it can be classified into: adjusting the speaking speed, content expansion, key points or difficult points, volume issues, etc.

[0111] Specifically, after classifying the interactive annotation and sending the at least one interactive annotation to the first terminal and each second terminal, the server captures an interactive audio with an interactive duration from the moment when the at least one interactive annotation is sent in the first report audio.

[0112] Then, the server performs speech recognition analysis on the interactive audio, wherein the speech recognition analysis may include, for example, speech rate analysis, semantic analysis, audio analysis, and the like. For example, when the interactive annotation sent by the server to the first terminal is classified as adjusting the speech rate, within the interactive audio of one interactive duration after the server sends the interactive annotation, the server performs speech rate analysis on the interactive audio, and uses the speech rate change of the interactive audio as the reporter's feedback information on the interactive annotation. When the interactive annotation sent by the server to the first terminal is classified as a volume problem, within the interactive audio of one interactive duration after the server sends the interactive annotation, the server performs audio analysis on the interactive audio, and uses the volume change of the interactive audio as the reporter's feedback information on the interactive annotation.

[0113] For example, when the interactive annotation sent by the server to the first terminal is classified as a key point or a difficult point, the server monitors the entire report audio after the interactive annotation from the first report audio, performs semantic analysis on the report audio, and matches the results of the semantic analysis with the key points or difficult points. The degree of matching between the results of the semantic analysis and the key points or difficult points is used as feedback information from the reporter regarding the interactive annotation.

[0114] It should be noted that if the server does not recognize the reporter's feedback information for any of the above interactive annotations, the feedback information will be considered empty.

[0115] 204: The server obtains a report evaluation for the reporter based on the reporter's electronic speech manuscript, the first report audio, the at least one interactive annotation, and the at least one feedback information.

[0116] In an embodiment of the present application, the server obtains the reporter's electronic speech manuscript, the first report audio, the above-mentioned at least one interactive annotation and the above-mentioned at least one feedback information, and performs intelligent analysis on the content of the entire meeting, for example, generating the reporter's report evaluation, performing data statistics, data mining, generating mind maps, etc.

[0117] Exemplarily, as shown in FIG8 , obtaining a report evaluation for the reporter based on the reporter's electronic speech manuscript, the first report audio, the at least one interactive annotation, and the at least one feedback information further includes the following steps:

[0118] 801: Based on the electronic speech manuscript and the first report audio, get the first evaluation.

[0119] In the embodiment of the present application, the first evaluation is used to indicate the reporter's completion quality for multiple preset nodes. The server can combine the content of the electronic speech manuscript and the reporter's first report audio to obtain the reporter's report completion quality for each preset node.

[0120] Exemplarily, obtaining a first evaluation based on the electronic speech manuscript and the first report audio also includes the following steps:

[0121] Sentence the electronic speech manuscript to obtain multiple first sentences;

[0122] Performing intent recognition on the plurality of first sentences to obtain a plurality of preset nodes corresponding to the plurality of first sentences;

[0123] Performing audio transcription on the first report audio to obtain an audio file, where the audio file includes a plurality of second sentences;

[0124] Matching the first sentence corresponding to each preset node with the plurality of second sentences respectively to obtain at least one second sentence corresponding to each preset node;

[0125] determining a fifth evaluation corresponding to each preset node based on the first sentence corresponding to each preset node, the at least one second sentence, and the type of each preset node;

[0126] The fifth evaluation corresponding to each preset node is weighted to obtain the first evaluation.

[0127] In this embodiment of the present application, the fifth evaluation is used to indicate the quality of the presenter's performance at each preset node. The server divides the electronic presentation into multiple preset nodes based on semantics, and determines the fifth evaluation corresponding to each preset node by matching the content in the first report audio with the preset node. The fifth evaluation is finally weighted and aggregated to obtain the first evaluation.

[0128] Specifically, the server first divides the text of the electronic speech into a plurality of first sentences according to the text end symbols thereof, and then performs intention recognition on the plurality of first sentences obtained above, wherein the intention recognition is used to analyze the functions of different sentence components of each first sentence.

[0129] Thus, multiple preset nodes corresponding to the above-mentioned multiple first sentences can be obtained. The multiple preset nodes include but are not limited to the following types: thinking nodes, interactive nodes, analysis nodes, and summary nodes. For example, thinking nodes are used to indicate the key points and difficulties in the electronic speech manuscript, that is, nodes that need to be left for participants to think about; interactive nodes are used to indicate content that requires interaction in the electronic speech manuscript; analysis nodes are used to indicate difficult points that require the presenter to answer questions for participants; and summary nodes are used to indicate places in the electronic speech manuscript that require annotations and summaries.

[0130] Furthermore, the first report audio is transcribed to obtain an audio file corresponding to the first report audio. The audio file is the text content of the audio translation, wherein the audio file includes multiple second sentences. Based on the position of each preset node, the first sentence corresponding to each preset node is matched with the multiple second sentences to obtain at least one second sentence corresponding to each preset node.

[0131] It should be noted that the reporter's report may add expanded content based on the content of the electronic speech manuscript. Before matching the first sentence corresponding to each preset node with multiple second sentences, it is also necessary to identify and analyze the content of each second sentence. The second sentences related to the content of the first sentence corresponding to each preset node before and after the position of the first sentence corresponding to each preset node are matched with the first sentence corresponding to each preset node, and finally at least one second sentence corresponding to each preset node is obtained.

[0132] Furthermore, an evaluation rule corresponding to each preset node is determined, and based on the evaluation rule, a fifth evaluation corresponding to each preset node is determined through the first sentence corresponding to each preset node, the at least one second sentence mentioned above, and the type of each preset node.

[0133] Exemplarily, the audio file further includes the start time and end time of each second sentence. Based on the first sentence corresponding to each preset node, the at least one second sentence, and the type of each preset node, determining the fifth evaluation corresponding to each preset node further includes the following steps:

[0134] If the type of the first preset node is a thinking node, the server obtains, from the audio file, an ending time of a first sentence corresponding to the first preset node and a starting time after the first sentence corresponding to the first preset node; obtains a duration between the starting time and the ending time; and determines, based on the duration, a fifth evaluation corresponding to the first preset node;

[0135] If the type of the first preset node is an analysis node, the server determines a first content that the reporter needs to analyze based on a first sentence corresponding to the first preset node; determines a second content that the reporter actually analyzes based on at least one second sentence corresponding to the first preset node; and determines a fifth evaluation corresponding to the first preset node based on the first content and the second content.

[0136] If the type of the first preset node is a summary node, the server determines the third content that the reporter needs to summarize based on the first sentence corresponding to the first preset node; determines the fourth content actually summarized by the reporter based on at least one second sentence corresponding to the first preset node; and determines the fifth evaluation corresponding to the first preset node based on the third content and the fourth content.

[0137] Exemplarily, if the type of the first preset node is an interactive node, the server determines the fifth content that the reporter needs to interact with based on the first sentence corresponding to the first preset node; determines the sixth content actually summarized by the reporter based on at least one second sentence corresponding to the first preset node; and determines the fifth evaluation corresponding to the first preset node based on the fifth content and the sixth content.

[0138] In the embodiment of the present application, the first preset node is any one of the plurality of preset nodes. The server may determine the completion quality of the reporter's report at each preset node based on the type of the first preset node and the corresponding different evaluation rules, i.e., determine the fifth evaluation corresponding to each preset node, and then aggregate each fifth evaluation to obtain the first evaluation.

[0139] Specifically, different weights are set for each type of preset node. For example, the weight of a thinking node is a1, the weight of an analysis node is a2, the weight of a summary node is a3, and the weight of an interaction node is a4. Then, the fifth evaluation corresponding to each thinking node is summed to obtain the total thinking node evaluation, the fifth evaluation corresponding to each analysis node is summed to obtain the total analysis node evaluation, the fifth evaluation corresponding to each summary node is summed to obtain the total summary node evaluation, and the fifth evaluation corresponding to each interaction node is summed to obtain the total interaction node evaluation.

[0140] Furthermore, the total thinking node evaluation is multiplied by the weight a1, the total analysis node evaluation is multiplied by the weight a2, the total summary node evaluation is multiplied by the weight a3, and the total interaction node evaluation is multiplied by the weight a4. Finally, the total thinking node evaluation, total analysis node evaluation, total summary node evaluation and total interaction node evaluation multiplied by their respective weights are summed to obtain the first evaluation.

[0141] It can be seen that by splitting the electronic speech into multiple first sentences and setting multiple preset nodes for the first sentences, the fifth evaluation corresponding to each preset node can be determined based on the above multiple preset nodes and audio files, and then the fifth evaluation is weighted and summarized to obtain the first evaluation, that is, the quality of the reporter's report completion for multiple preset nodes, which in turn helps the reporter adjust the content of the report and improves the efficiency and accuracy of the meeting report.

[0142] 802: Analyze the first report audio and obtain a second evaluation.

[0143] In the embodiment of the present application, the second evaluation is used to indicate the quality of the reporter's entire report. The server analyzes the entire meeting report process based on the reporter's first report audio and obtains an evaluation of the overall dimension of the report, namely the second evaluation.

[0144] Specifically, the server presets a pause time threshold for determining whether the reporter is stuttering based on the pause time between each sentence in the report audio. For example, the pause time threshold is 3 seconds. The server then performs audio analysis on the first report audio to obtain audio files of multiple pause points, and uses the time of each pause point in the audio file as the pause time of each pause point.

[0145] If the pause time at the first pause point is greater than the pause time threshold, it indicates that the reporter has stalled at the first pause point, and the number of stalls of the reporter is increased by one; otherwise, the first pause point is a normal pause, and the pause time of the next pause point is detected. The first pause point is any of the multiple pause points mentioned above.

[0146] Furthermore, the server obtains the total reporting time of the first reporting audio, and uses the ratio of the reporter's number of pauses to the total reporting time of the first reporting audio as the reporter's fluency during the reporting process. The reporter's fluency during the reporting process and the number of pauses during the reporting process are summarized to obtain a second evaluation.

[0147] 803: Based on the at least one interactive annotation, determine a third evaluation corresponding to the at least one interactive annotation.

[0148] In the embodiment of the present application, the third evaluation is used to represent the evaluation of the reporter's report by the participant. The server analyzes at least one interactive annotation obtained from the plurality of second terminals to obtain the evaluation of each participant on the reporter's report.

[0149] Exemplarily, based on the at least one interactive annotation, determining a third evaluation corresponding to the at least one interactive annotation may further include the following steps:

[0150] For each interactive annotation, the server classifies each interactive annotation and obtains the category corresponding to each interactive annotation;

[0151] The server obtains the identity information of the participant who sent each interactive annotation;

[0152] Based on the identity information of the participant of each interactive annotation, the server determines the weight of each interactive annotation under each interactive annotation category;

[0153] The server groups the categories corresponding to at least one interactive annotation to obtain at least one category group;

[0154] Based on the weight of each interactive annotation under each interactive annotation category, the server determines the weight corresponding to each category group;

[0155] The server determines a sixth evaluation corresponding to each category group based on the weight corresponding to each category group and the score corresponding to each category group;

[0156] The server obtains a third evaluation corresponding to at least one interactive annotation based on the sixth evaluation corresponding to each category group.

[0157] Specifically, the server classifies each interactive comment using, for example, a text classification algorithm in natural language processing, to obtain a category corresponding to each interactive comment. For example, the categories corresponding to the interactive comment include, but are not limited to, comments, questions, praise, negative, and so on.

[0158] Furthermore, the server obtains the identity information of the attendees of each interactive annotation, and assigns corresponding weights to the number of categories corresponding to each interactive annotation based on the identity information of the attendees of each interactive annotation. For example, when the first interactive annotation is praise, if the attendee of the first interactive annotation is a leader, when counting the interactive annotations in the praise category, the number of the first interactive annotation is multiplied by the weight a1, that is, the number of interactive annotations in the praise category plus a1, where a1 is greater than 1; otherwise, the number of interactive annotations in the praise category is increased by one. The first interactive annotation is any one of the at least one interactive annotation mentioned above. When the first interactive annotation is a comment, question, or negative, the method of assigning corresponding weights to the number of categories corresponding to each interactive annotation is the same as the above, and will not be described here.

[0159] Furthermore, the server obtains the page dwell time of the participants of each interactive annotation before sending the interactive annotation, that is, the attention level of the participants of each interactive annotation. Based on the attention level of the participants of each interactive annotation, the weight of the number of categories corresponding to each interactive annotation is adjusted. For example, if the page dwell time of the participants of the second interactive annotation is less than the preset page dwell time, the weight corresponding to the number of categories of the second interactive annotation is adjusted to a2, where a2 is less than a1; otherwise, the weight corresponding to the number of categories of the second interactive annotation is a1. The second interactive annotation is any one of the at least one interactive annotation mentioned above.

[0160] Furthermore, the categories corresponding to the at least one interactive annotation are grouped to obtain at least one category group, for example, the category group includes: positive, negative, and neutral. The weighted count of the categories of each interactive annotation is calculated using the above steps, and the weight corresponding to each category group is calculated by adding the categories of the at least one interactive annotation in the first category group. The first category group is any one of the at least one category group.

[0161] The server then assigns a corresponding score to each category group and multiplies the score by the weight to obtain a sixth evaluation for each category group. The sixth evaluations for each category group are then aggregated to obtain a third evaluation for the at least one interactive annotation.

[0162] It can be seen that by classifying each interactive annotation, based on the participant information and attention of each interactive annotation, the weight of each interactive annotation under each interactive annotation category can be determined, thereby determining the weight corresponding to each interaction corresponding to the category group, and then according to the corresponding score of each category group, determining the sixth evaluation corresponding to each category group, and summarizing to obtain the third evaluation, so that the reporter can understand each participant's evaluation of the reporter's report through the evaluation system, and adjust the report content, thereby improving the reporting efficiency and accuracy of the meeting.

[0163] 804: Obtain a fourth evaluation based on the at least one interactive annotation and the at least one feedback information.

[0164] In the embodiment of the present application, the fourth evaluation is used to indicate the reporter's processing efficiency of the interactive annotations of the participants.

[0165] Exemplarily, obtaining a fourth evaluation based on the at least one interactive comment and the at least one feedback information may further include the following steps:

[0166] The server performs intent recognition on each interactive annotation to obtain the purpose of each interactive annotation;

[0167] Based on the feedback information corresponding to each interactive annotation, the server determines the reporter's probability of completing the purpose of each interactive annotation;

[0168] Based on the completion probability of the purpose of each interactive annotation, the server determines the evaluation corresponding to each interactive annotation;

[0169] The server obtains a fourth evaluation based on the evaluation corresponding to each interactive annotation and the number of the at least one interactive annotation.

[0170] Specifically, the server performs intent recognition on each interactive annotation, determines the purpose of each interactive annotation, and predicts the predicted feedback information for each interactive annotation. The server then performs semantic recognition on the feedback information corresponding to each interactive annotation, determining the corresponding feedback information for each interactive annotation. The server then compares the feedback information corresponding to each interactive annotation with the predicted feedback information for each interactive annotation, and obtains the number of interactive annotations that successfully matched.

[0171] Furthermore, the ratio of the number of feedback messages for the successfully matched interactive annotations to the total number of feedback messages for the interactive annotations is used as the reporter's probability of achieving the purpose of each interactive annotation. The probability of achieving the purpose of each interactive annotation is multiplied by the preset score to obtain an evaluation corresponding to each interactive annotation. The evaluation corresponding to each interactive annotation is multiplied by the number of the at least one interactive annotation to obtain a fourth evaluation.

[0172] It can be seen that through the feedback information corresponding to each interactive annotation, the reporter's feedback on each interactive annotation can be determined, thereby obtaining a fourth evaluation based on the reporter, so that the reporter can obtain the completion status of the feedback on the participant's interactive annotation, thereby adjusting the report content and improving the accuracy of the meeting report.

[0173] 805: Based on the first evaluation, the second evaluation, the third evaluation, and the fourth evaluation, obtain a report evaluation for the reporter.

[0174] In the embodiment of the present application, after obtaining the first evaluation, the second evaluation, the third evaluation, and the fourth evaluation, the server summarizes each evaluation to obtain a report evaluation and a report briefing for the reporter.

[0175] Furthermore, the server segments the presenter's report based on its content and theme, extracts keywords from each segment, and obtains keywords for each chapter. By mining the relationships between each keyword, it then obtains the chapter theme and keywords for each segment. A mind map is generated based on the connections between the report content and theme, and between each chapter theme and keywords.

[0176] 205: The server sends the evaluation report to the first terminal.

[0177] In an embodiment of the present application, the server sends the report evaluation, report briefing and mind map to the first terminal. After receiving the report evaluation, report briefing and mind map, the first terminal displays the report evaluation, report briefing and mind map on the display interface of the first terminal.

[0178] It can be seen that the server determines the completion quality of the meeting report in the first report audio through the content of the electronic speech manuscript, and obtains the first evaluation. The second evaluation is obtained through voice analysis of the first report audio. The third evaluation is obtained by monitoring the interactive annotations of the participants on the reporter's report content. The fourth evaluation is obtained through the feedback information of the reporter on the interactive annotations of the participants. The first evaluation, second evaluation, third evaluation and fourth evaluation are then analyzed and summarized to obtain the report evaluation for the reporter, so that the reporter can obtain reference information from the report evaluation and adjust the content of the meeting report, thereby improving the efficiency and accuracy of the report.

[0179] Refer to Figure 9, which is a block diagram of the functional units of a paperless real-time evaluation report processing device provided in an embodiment of the present application. As shown in Figure 9, the paperless real-time evaluation report processing device 900 includes an acquisition unit 901 and a processing unit 902.

[0180] An acquiring unit 901 is configured to acquire, from a first terminal, an electronic speech manuscript of a reporter and a first audio report of the electronic speech manuscript, wherein the first audio report is the audio of the reporter during the entire reporting process;

[0181] The processing unit 902 is configured to obtain, from at least one second terminal among the plurality of second terminals, at least one interactive annotation of a participant on a report of a reporter;

[0182] Obtain at least one piece of feedback information from the reporter regarding the at least one interactive comment;

[0183] Obtaining a report evaluation of the reporter based on the reporter's electronic speech manuscript, the first report audio, the at least one interactive annotation, and the at least one feedback information;

[0184] Sending a report evaluation to the first terminal.

[0185] In one embodiment of the present application, in obtaining a report evaluation for the reporter based on the reporter's electronic speech manuscript, the first report audio, the at least one interactive annotation, and the at least one feedback information, the processing unit 902 is specifically configured to:

[0186] Obtaining a first evaluation based on the electronic speech manuscript and the first report audio, wherein the first evaluation is used to indicate the reporter's completion quality for a plurality of preset nodes;

[0187] Analyze the first report audio to obtain a second evaluation, wherein the second evaluation is used to indicate the reporting quality of the reporter's entire report;

[0188] Based on the at least one interactive annotation, determining a third evaluation corresponding to the at least one interactive annotation, wherein the third evaluation is used to represent the evaluation of the participant on the reporter's report;

[0189] Based on the at least one interactive annotation and the at least one piece of feedback information, a fourth evaluation is obtained, wherein the fourth evaluation is used to indicate the efficiency of the reporter in processing the interactive annotation of the participant;

[0190] Based on the first evaluation, the second evaluation, the third evaluation, and the fourth evaluation, a report evaluation of the reporter is obtained.

[0191] In one embodiment of the present application, in obtaining the first evaluation based on the electronic speech manuscript and the first report audio, the processing unit 902 is specifically configured to:

[0192] Sentence the electronic speech manuscript to obtain multiple first sentences;

[0193] Performing intent recognition on the plurality of first sentences to obtain a plurality of preset nodes corresponding to the plurality of first sentences;

[0194] Performing audio transcription on the first report audio to obtain an audio file, where the audio file includes a plurality of second sentences;

[0195] Matching the first sentence corresponding to each preset node with the plurality of second sentences respectively to obtain at least one second sentence corresponding to each preset node;

[0196] Determining a fifth evaluation corresponding to each preset node based on the first sentence corresponding to each preset node, the at least one second sentence, and the type of each preset node, wherein the fifth evaluation is used to indicate the reporter's completion quality at each preset node;

[0197] The fifth evaluation corresponding to each preset node is weighted to obtain the first evaluation.

[0198] In one embodiment of the present application, the audio file further includes a start time and an end time of each second sentence; in determining the fifth evaluation corresponding to each preset node based on the first sentence corresponding to each preset node, the at least one second sentence, and the type of each preset node, the processing unit 902 is specifically configured to:

[0199] If the type of the first preset node is a thinking node, obtaining from the audio file an end time of a first sentence corresponding to the first preset node and a start time located after the first sentence corresponding to the first preset node; obtaining a duration between the start time and the end time; and determining a fifth evaluation corresponding to the first preset node based on the duration, wherein the first preset node is any one of the plurality of preset nodes;

[0200] If the type of the first preset node is an analysis node, determining a first content that the reporter needs to analyze based on a first sentence corresponding to the first preset node; determining a second content that the reporter actually analyzes based on at least one second sentence corresponding to the first preset node; and determining a fifth evaluation corresponding to the first preset node based on the first content and the second content.

[0201] If the type of the first preset node is a summary node, determine the third content that the reporter needs to summarize based on the first sentence corresponding to the first preset node; determine the fourth content actually summarized by the reporter based on at least one second sentence corresponding to the first preset node; and determine the fifth evaluation corresponding to the first preset node based on the third content and the fourth content.

[0202] In one embodiment of the present application, in determining the third evaluation corresponding to the at least one interactive annotation based on the at least one interactive annotation, the processing unit 902 is specifically configured to:

[0203] For each interactive annotation, classify each interactive annotation to obtain the category corresponding to each interactive annotation;

[0204] Get the identity information of the participant who sent each interactive annotation;

[0205] Based on the identity information of the participant of each interactive annotation, determine the weight of each interactive annotation under each interactive annotation category;

[0206] Grouping the categories corresponding to the at least one interactive annotation to obtain at least one category group;

[0207] Based on the weight of each interactive annotation under the category of each interactive annotation, determine the weight corresponding to each category group;

[0208] determining a sixth evaluation corresponding to each category group based on the weight corresponding to each category group and the score corresponding to each category group;

[0209] Based on the sixth evaluation corresponding to each category group, a third evaluation corresponding to the at least one interactive annotation is obtained.

[0210] In one embodiment of the present application, in obtaining the fourth evaluation based on the at least one interactive annotation and the at least one feedback information, the processing unit 902 is specifically configured to:

[0211] Perform intent recognition on each interactive annotation to obtain the purpose of each interactive annotation;

[0212] Based on the feedback information corresponding to each interactive annotation, determine the reporter's completion probability of the purpose of each interactive annotation;

[0213] Determine the evaluation corresponding to each interactive annotation based on the completion probability of the purpose of each interactive annotation;

[0214] A fourth evaluation is obtained based on the evaluation corresponding to each interactive annotation and the number of the at least one interactive annotation.

[0215] In one embodiment of the present application, the processing unit 902 is further configured to:

[0216] Obtaining a synchronization request from the first terminal, wherein the synchronization request is used to request provision of a speech script synchronization service for the reporter;

[0217] In response to the synchronization request, obtaining a second reporting audio of the reporter at the current moment from the first terminal;

[0218] Perform semantic analysis on the second report audio to obtain the first report position of the reporter at the current moment;

[0219] Perform speech speed analysis on the second report audio to obtain the current speaking speed of the reporter;

[0220] Based on the reporter's current speaking speed, the reporter's first reporting position, and the electronic speech manuscript, predict the reporter's second reporting position at the next moment;

[0221] Determining the number of words the reporter needs to report from the current moment to the next moment based on the second reporting position and the first reporting position;

[0222] Determine the playback speed of the electronic speech based on the number of words and the time between the current moment and the next moment;

[0223] The playing speed and the first reporting position are synchronized to the first terminal and the plurality of second terminals, so that the first terminal and the plurality of second terminals all play the electronic speech manuscript from the first reporting position and according to the playing speed at the current moment.

[0224] Referring to Figure 10 , Figure 10 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. As shown in Figure 10 , electronic device 1000 includes a transceiver 1001, a processor 1002, and a memory 1003. These are connected via a bus 1004. Memory 1003 is used to store computer programs and data and can transmit data stored in memory 1003 to processor 1002. Electronic device 1000 may be the paperless real-time evaluation report processing device 900 shown in Figure 9 above.

[0225] The processor 1002 is configured to read the computer program in the memory 1003 and perform the following operations:

[0226] Obtaining from the first terminal an electronic speech manuscript of the reporter and a first report audio of the electronic speech manuscript, wherein the first report audio is the audio of the reporter during the entire report process;

[0227] obtaining, from at least one second terminal among the plurality of second terminals, at least one interactive annotation of a participant on the report of the reporter;

[0228] Obtain at least one piece of feedback information from the reporter regarding the at least one interactive comment;

[0229] Obtaining a report evaluation of the reporter based on the reporter's electronic speech manuscript, the first report audio, the at least one interactive annotation, and the at least one feedback information;

[0230] Sending a report evaluation to the first terminal.

[0231] In one embodiment of the present application, in obtaining a report evaluation for the reporter based on the reporter's electronic speech manuscript, the first report audio, the at least one interactive annotation, and the at least one feedback information, the processor 1002 is specifically configured to perform the following operations:

[0232] Obtaining a first evaluation based on the electronic speech manuscript and the first report audio, wherein the first evaluation is used to indicate the reporter's completion quality for a plurality of preset nodes;

[0233] Analyze the first report audio to obtain a second evaluation, wherein the second evaluation is used to indicate the reporting quality of the reporter's entire report;

[0234] Based on the at least one interactive annotation, determining a third evaluation corresponding to the at least one interactive annotation, wherein the third evaluation is used to represent the evaluation of the participant on the reporter's report;

[0235] Based on the at least one interactive annotation and the at least one piece of feedback information, a fourth evaluation is obtained, wherein the fourth evaluation is used to indicate the efficiency of the reporter in processing the interactive annotation of the participant;

[0236] Based on the first evaluation, the second evaluation, the third evaluation, and the fourth evaluation, a report evaluation of the reporter is obtained.

[0237] In one embodiment of the present application, in obtaining the first evaluation based on the electronic speech manuscript and the first report audio, the processor 1002 is specifically configured to perform the following operations:

[0238] Sentence the electronic speech manuscript to obtain multiple first sentences;

[0239] Performing intent recognition on the plurality of first sentences to obtain a plurality of preset nodes corresponding to the plurality of first sentences;

[0240] Performing audio transcription on the first report audio to obtain an audio file, where the audio file includes a plurality of second sentences;

[0241] Matching the first sentence corresponding to each preset node with the plurality of second sentences respectively to obtain at least one second sentence corresponding to each preset node;

[0242] Determining a fifth evaluation corresponding to each preset node based on the first sentence corresponding to each preset node, the at least one second sentence, and the type of each preset node, wherein the fifth evaluation is used to indicate the reporter's completion quality at each preset node;

[0243] The fifth evaluation corresponding to each preset node is weighted to obtain the first evaluation.

[0244] In one embodiment of the present application, the audio file further includes a start time and an end time of each second sentence; in determining the fifth evaluation corresponding to each preset node based on the first sentence corresponding to each preset node, the at least one second sentence, and the type of each preset node, the processor 1002 is specifically configured to perform the following operations:

[0245] If the type of the first preset node is a thinking node, obtaining from the audio file an end time of a first sentence corresponding to the first preset node and a start time located after the first sentence corresponding to the first preset node; obtaining a duration between the start time and the end time; and determining a fifth evaluation corresponding to the first preset node based on the duration, wherein the first preset node is any one of the plurality of preset nodes;

[0246] If the type of the first preset node is an analysis node, determining a first content that the reporter needs to analyze based on a first sentence corresponding to the first preset node; determining a second content that the reporter actually analyzes based on at least one second sentence corresponding to the first preset node; and determining a fifth evaluation corresponding to the first preset node based on the first content and the second content.

[0247] If the type of the first preset node is a summary node, determine the third content that the reporter needs to summarize based on the first sentence corresponding to the first preset node; determine the fourth content actually summarized by the reporter based on at least one second sentence corresponding to the first preset node; and determine the fifth evaluation corresponding to the first preset node based on the third content and the fourth content.

[0248] In one embodiment of the present application, in determining the third evaluation corresponding to the at least one interactive annotation based on the at least one interactive annotation, the processor 1002 is specifically configured to perform the following operations:

[0249] For each interactive annotation, classify each interactive annotation to obtain the category corresponding to each interactive annotation;

[0250] Get the identity information of the participant who sent each interactive annotation;

[0251] Based on the identity information of the participant of each interactive annotation, determine the weight of each interactive annotation under each interactive annotation category;

[0252] Grouping the categories corresponding to the at least one interactive annotation to obtain at least one category group;

[0253] Based on the weight of each interactive annotation under the category of each interactive annotation, determine the weight corresponding to each category group;

[0254] determining a sixth evaluation corresponding to each category group based on the weight corresponding to each category group and the score corresponding to each category group;

[0255] Based on the sixth evaluation corresponding to each category group, a third evaluation corresponding to the at least one interactive annotation is obtained.

[0256] In one embodiment of the present application, in obtaining the fourth evaluation based on the at least one interactive annotation and the at least one feedback information, the processor 1002 is specifically configured to perform the following operations:

[0257] Perform intent recognition on each interactive annotation to obtain the purpose of each interactive annotation;

[0258] Based on the feedback information corresponding to each interactive annotation, determine the reporter's completion probability of the purpose of each interactive annotation;

[0259] Determine the evaluation corresponding to each interactive annotation based on the completion probability of the purpose of each interactive annotation;

[0260] A fourth evaluation is obtained based on the evaluation corresponding to each interactive annotation and the number of the at least one interactive annotation.

[0261] In one embodiment of the present application, the processor 1002 is further configured to perform the following operations:

[0262] Obtaining a synchronization request from the first terminal, wherein the synchronization request is used to request provision of a speech script synchronization service for the reporter;

[0263] In response to the synchronization request, obtaining a second reporting audio of the reporter at the current moment from the first terminal;

[0264] Perform semantic analysis on the second report audio to obtain the first report position of the reporter at the current moment;

[0265] Perform speech speed analysis on the second report audio to obtain the current speaking speed of the reporter;

[0266] Based on the reporter's current speaking speed, the reporter's first reporting position, and the electronic speech manuscript, predict the reporter's second reporting position at the next moment;

[0267] Determining the number of words the reporter needs to report from the current moment to the next moment based on the second reporting position and the first reporting position;

[0268] Determine the playback speed of the electronic speech based on the number of words and the time between the current moment and the next moment;

[0269] The playing speed and the first reporting position are synchronized to the first terminal and the plurality of second terminals, so that the first terminal and the plurality of second terminals all play the electronic speech manuscript from the first reporting position and according to the playing speed at the current moment.

[0270] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. The computer program is executed by a processor to implement part or all of the steps of any paperless real-time evaluation report processing method described in the above method embodiments.

[0271] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute part or all of the steps of any paperless real-time evaluation report processing method recorded in the above method embodiments.

[0272] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.

[0273] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0274] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0275] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0276] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of software program modules.

[0277] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0278] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0279] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for those skilled in the art, according to the idea of ​​the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A paperless real-time evaluation and reporting processing method, characterized in that, the method is applied to a paperless real-time evaluation and reporting processing device, the paperless real-time evaluation and reporting processing device is located in a paperless real-time evaluation and reporting processing system, the paperless real-time evaluation and reporting processing system further includes a first terminal and a plurality of second terminals, wherein the first terminal is the terminal of the reporter, the plurality of second terminals are the terminals of a plurality of participants, and the plurality of second terminals correspond to the plurality of participants one by one; the method includes: obtaining the electronic speech script of the reporter and a first reporting audio for the electronic speech script from the first terminal, wherein the first reporting audio is the audio of the reporter during the entire reporting process; obtaining at least one interactive annotation of the participant for the reporter's report from at least one of the plurality of second terminals; obtaining at least one feedback message of the reporter for the at least one interactive annotation; obtaining a report evaluation for the reporter's report based on the electronic speech script of the reporter, the first reporting audio, the at least one interactive annotation, and the at least one feedback message; sending the report evaluation to the first terminal.

2. The method according to claim 1, characterized in that, the obtaining a report evaluation for the reporter's report based on the electronic speech script of the reporter, the first reporting audio, the at least one interactive annotation, and the at least one feedback message includes: obtaining a first evaluation based on the electronic speech script and the first reporting audio, wherein the first evaluation is used to indicate the completion quality of the reporter for a plurality of preset nodes; analyzing the first reporting audio to obtain a second evaluation, wherein the second evaluation is used to indicate the reporting quality of the reporter for the entire report; determining a third evaluation corresponding to the at least one interactive annotation based on the at least one interactive annotation, wherein the third evaluation is used to characterize the evaluation of the participant for the reporter's report; obtaining a fourth evaluation based on the at least one interactive annotation and the at least one feedback message, wherein the fourth evaluation is used to indicate the processing efficiency of the reporter for the interactive annotation of the participant; obtaining a report evaluation for the reporter's report based on the first evaluation, the second evaluation, the third evaluation, and the fourth evaluation.

3. The method according to claim 2, characterized in that, the obtaining a first evaluation based on the electronic speech script and the first reporting audio includes: clause-dividing the electronic speech script to obtain a plurality of first sentences; performing intention recognition on the plurality of first sentences to obtain a plurality of preset nodes corresponding to the plurality of first sentences; performing audio transcription on the first reporting audio to obtain an audio file, the audio file including a plurality of second sentences; matching each first sentence corresponding to a preset node with the plurality of second sentences respectively to obtain at least one second sentence corresponding to each preset node; Determine a fifth evaluation corresponding to each preset node based on the first sentence corresponding to each preset node, the at least one second sentence, and the type of each preset node, where the fifth evaluation is used to indicate the completion quality of the reporter at each preset node; Weight the fifth evaluation corresponding to each preset node to obtain the first evaluation.

4. The method according to claim 3, wherein, the audio file further includes the start time and end time of each second sentence; the determining a fifth evaluation corresponding to each preset node based on the first sentence corresponding to each preset node, the at least one second sentence, and the type of each preset node includes: If the type of the first preset node is a thinking node, obtain the end time of the first sentence corresponding to the first preset node from the audio file, and the start time after the first sentence corresponding to the first preset node; obtain the duration between the start time and the end time; determine the fifth evaluation corresponding to the first preset node according to the duration, where the first preset node is any one of the multiple preset nodes; If the type of the first preset node is an analysis node, determine the first content that the reporter needs to analyze according to the first sentence corresponding to the first preset node; determine the second content that the reporter actually analyzes according to the at least one second sentence corresponding to the first preset node; determine the fifth evaluation corresponding to the first preset node according to the first content and the second content; If the type of the first preset node is a summary node, determine the third content that the reporter needs to summarize according to the first sentence corresponding to the first preset node; determine the fourth content that the reporter actually summarizes according to the at least one second sentence corresponding to the first preset node; determine the fifth evaluation corresponding to the first preset node according to the third content and the fourth content.

5. The method according to any one of claims 2-4, wherein, the determining a third evaluation corresponding to the at least one interactive annotation based on the at least one interactive annotation includes: For each interactive annotation, classify each interactive annotation to obtain the category corresponding to each interactive annotation; Obtain the identity information of the attendee who sent each interactive annotation; Based on the identity information of the attendee of each interactive annotation, determine the weight of each interactive annotation under the category of each interactive annotation; Group the categories corresponding to the at least one interactive annotation to obtain at least one category group; Based on the weight of each interactive annotation under the category of each interactive annotation, determine the weight corresponding to each category group; Based on the weight corresponding to each category group and the score corresponding to each category group, determine the sixth evaluation corresponding to each category group; Based on the sixth evaluation corresponding to each category group, obtain the third evaluation corresponding to the at least one interactive annotation.

6. The method according to any one of claims 2-5, wherein, the obtaining a fourth evaluation based on the at least one interactive annotation and the at least one feedback message includes: Perform intent recognition on each interactive annotation to obtain the purpose of each interactive annotation; Based on the feedback information corresponding to each interactive annotation, determine the completion probability of the purpose of each interactive annotation by the reporter; Based on the completion probability of the purpose of each interactive annotation, determine the evaluation corresponding to each interactive annotation; Based on the evaluation corresponding to each interactive annotation and the number of the at least one interactive annotation, obtain a fourth evaluation.

7. The method according to any one of claims 1-6, characterized in that, the method further includes: Obtain a synchronization request from the first terminal, where the synchronization request is used to request to provide a speech script synchronization service for the reporter; In response to the synchronization request, obtain the second reporting audio of the reporter at the current moment from the first terminal; Perform semantic analysis on the second reporting audio to obtain the first reporting position of the reporter at the current moment; Perform speech rate analysis on the second reporting audio to obtain the speech rate of the reporter at the current moment; Based on the speech rate of the reporter at the current moment, the first reporting position of the reporter, and the electronic speech script, predict the second reporting position of the reporter at the next moment; Based on the second reporting position and the first reporting position, determine the reporter from the current moment to the next moment The number of words to be reported; Based on the number of words and the duration between the current moment and the next moment, determine the playback speed of the electronic speech script; Synchronize the playback speed and the first reporting position to the first terminal and the multiple second terminals, so that the first terminal and the multiple second terminals both start from the first reporting position at the current moment and play the electronic speech script at the playback speed.

8. A paperless real-time evaluation reporting processing device, characterized in that, the paperless real-time evaluation reporting processing device includes an acquisition unit and a processing unit; The acquisition unit is used to obtain the electronic speech script of the reporter and the first reporting audio for the electronic speech script from the first terminal, where the first reporting audio is the audio of the reporter during the entire reporting process; The processing unit is used to obtain at least one interactive annotation for the reporter's report from at least one of the multiple second terminals; Obtain at least one feedback information of the reporter for the at least one interactive annotation; Based on the electronic speech script of the reporter, the first reporting audio, the at least one interactive annotation, and the at least one feedback information, obtain a report evaluation for the reporter; Send the report evaluation to the first terminal.

9. An electronic device, characterized in that, includes: A processor and a memory, the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Artificial intelligence-based lecturer dynamic evaluation method and device and computer equiment

    CN112232166A

  • Method and equipment for evaluating logic structure of speech draft

    CN113361275A

  • Video conference monitoring system

    CN117156122A