Session data processing method and electronic device
By acquiring and analyzing session data in real time and using models to adjust the meeting schedule, the problem of inappropriate speaking by speakers in project reporting meetings has been solved, achieving more efficient and controllable meeting management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LENOVO (BEIJING) LTD
- Filing Date
- 2026-03-24
- Publication Date
- 2026-07-31
AI Technical Summary
In project presentation meetings, existing technologies struggle to effectively control the content and duration of each speaker's presentation, leading to reduced meeting efficiency and predictability.
By acquiring conversation data in real time and using models to analyze the content and duration of speeches, the progress of the conversation is dynamically adjusted, including generating prompts to guide speakers to supplement or end their speeches, ensuring that the coverage and duration of the speech content meet the predetermined requirements.
This improved the efficiency and controllability of the meeting, ensuring that each speaker conveyed the necessary information within a reasonable timeframe, and enhancing the comprehensiveness and effectiveness of the meeting.
Smart Images

Figure CN122496493A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a voice data processing method and electronic device. Background Technology
[0002] Project reporting meetings are a common form of structured communication. Typically, the facilitator asks each project leader to report on specific key points, such as the project's main pain points, proposed solutions, implementation methods, and progress. To improve meeting efficiency, the facilitator sets a fixed reporting time for each project, such as 10 minutes. However, in actual meetings, differences in presenters' speaking abilities, preparation levels, and logical flow can lead to incomplete or off-topic presentations, and the presentation duration is difficult to control, thus reducing the overall efficiency and predictability of the meeting.
[0003] Therefore, without changing the format of the meeting, a method is needed to dynamically adjust the content and pace of the presentation. Summary of the Invention
[0004] This disclosure provides a session data processing method and an electronic device to at least solve the above-mentioned technical problems existing in the prior art.
[0005] According to a first aspect of this disclosure, a session data processing method is provided, the method comprising:
[0006] During the conversation, conversation data is obtained; the conversation data includes at least the content and duration of the current speaker's speech. The session data is parsed based on the model; The processing logic is determined based on the parsing results of the session data, and the processing logic is executed to control the progress of the session.
[0007] According to a second aspect of this disclosure, a session data processing apparatus is provided, the apparatus comprising: The acquisition module is used to acquire session data during the session; the session data includes at least the content and duration of the current speaker's speech. The processing module is used to parse the session data based on the model; determine the processing logic based on the parsing result of the session data; and execute the processing logic to control the progress of the session.
[0008] According to a third aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the session data processing method described in this disclosure.
[0009] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the session data processing method described in this disclosure.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which: In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.
[0012] Figure 1 A flowchart illustrating a session data processing method provided in an embodiment of this disclosure; Figure 2 A flowchart illustrating a meeting method based on content coverage and time prediction provided in an embodiment of this disclosure; Figure 3 A flowchart illustrating a session data processing method provided in an embodiment of this disclosure; Figure 4 A flowchart illustrating a session data processing method provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of the structure of a session data processing apparatus provided in an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0013] To make the objectives, features, and advantages of this disclosure more apparent and understandable, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0014] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0015] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in an order other than that illustrated or described herein.
[0016] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in this disclosure is for the purpose of describing embodiments of this disclosure only and is not intended to be limiting of this disclosure.
[0017] It should be understood that in the various embodiments of this disclosure, the sequence number of each implementation process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure.
[0018] Figure 1 This is a flowchart illustrating a session data processing method provided in an embodiment of the present disclosure, as shown below. Figure 1 As shown, the method includes: Step 101: During the conversation, obtain conversation data; the conversation data includes at least the content and duration of the current speaker's speech; Step 102: Parse the session data based on the model; Step 103: Determine the processing logic based on the parsing results of the session data, and execute the processing logic to control the progress of the session.
[0019] Here, the method can be applied to electronic devices such as servers, computers, mobile phones, smartphones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), wearable devices (such as smart bracelets, smartwatches, etc.).
[0020] The session data is collected in real time during the meeting and may include at least the content and duration of the current speaker's remarks. The meeting may involve multiple participants, who are categorized as speakers and other participants in each round of speaking.
[0021] The model can be a specific model that is pre-designed and invoked by an electronic device. This model can be used to parse and analyze the collected conversation data. The model can be based on natural language processing technology to extract key points and design elements in the speech, and to identify the speaker's performance, etc.
[0022] Based on the analysis results of the session data by the model, the electronic device can determine the corresponding processing logic. The analysis results can represent the results related to the content of the speech and the results related to the duration of the speech. By judging whether the results related to the content of the speech and the results related to the duration of the speech meet the corresponding conditions, the corresponding processing logic is determined. The processing logic is used to adjust the progress of the session.
[0023] In this way, electronic devices improve the efficiency and controllability of the session through real-time data analysis, ensuring that the session can proceed according to the predetermined goals and time, and realizing automatic adjustment of the session progress.
[0024] In some embodiments, the processing logic is determined based on the parsing result of the session data, including: In response to the session data satisfying a first condition, a first processing logic is determined, which is used to end the current speech; The first condition indicates that the content coverage of the conversation data reaches a first threshold and the speaking duration does not reach a second threshold.
[0025] Here, the first condition includes two aspects: the coverage of the speech content and the duration of the speech.
[0026] The coverage of the speech content refers to the comprehensiveness and sufficiency of the key points or critical information involved in the speech; the coverage of the speech content reaching the first threshold indicates that the current speaker has covered the required key points or critical information in the speech.
[0027] The speaking duration represents the time currently used by the speaker. The speaker not reaching the second threshold means that the speaker's speaking time has not yet reached the set time limit and the speaker can continue speaking.
[0028] It should be noted that the first and second thresholds are thresholds set based on application requirements, and there are no restrictions on their values.
[0029] Thus, the dynamic adjustment mechanism based on conversation data analysis ensures that speakers convey the necessary information within a reasonable timeframe, thereby improving meeting efficiency.
[0030] In one example, the host or organizer pre-sets the essential content structure for the session, including the following key points: pain points, solutions, implementation paths, and risks; and sets time limits for each point of discussion. For example, the first limit is a 30-minute limit for the overall session; the second limit is a 10-minute limit for each of the key points, such as pain points, solutions, implementation paths, and risks. Of course, the time limit for each key point can be different; this is just one example, and the duration is not limited; the third limit is a combination of the first and second limits, i.e., a 30-minute limit for the overall session, and a separate time limit for each key point.
[0031] During the conversation, conversation data is collected in real time, including the current speaker's speech content and duration; the conversation data is then parsed, and the parsing results are assumed to include: The current presentation has covered some of the key points of the project, such as pain points, solutions, implementation paths, and risks, which meets the preset first threshold. The speaking time limit was 20 minutes, but the meeting stipulated that the maximum time limit for each speaker was 30 minutes, meaning that the speaking time did not reach the second threshold.
[0032] In the current situation, if the analysis results indicate that the speech has covered the required information, the current speech can end. For example, the project leader can be notified through the moderator or automatic prompts that their speech has covered the required information and they can proceed to the next round of the agenda. Alternatively, the speaking privileges of the previous speaker can be closed, and the speaking privileges of the next speaker can be activated to proceed to the next round of speeches.
[0033] In some embodiments, the processing logic is determined based on the parsing result of the session data, including: In response to the session data satisfying the second condition, a second processing logic is determined, the second processing logic including generating a first prompt message based on the session data and outputting the first prompt message to the current speaker; The second condition indicates that the content coverage of the conversation data does not reach the first threshold and the speaking duration does not reach the second threshold. The first prompt message is used to guide the current speaker to speak based on the content of the first prompt message in order to improve the coverage of the speech content.
[0034] Here, the second condition includes two aspects: the coverage of the speech content and the duration of the speech.
[0035] The coverage of the speech content refers to the comprehensiveness and sufficiency of the key points or information involved in the speech; the coverage of the speech content not reaching the first threshold indicates that the current speaker has not fully covered all the necessary topics or information points in the speech.
[0036] The speaking duration represents the time currently used by the speaker. The speaker not reaching the second threshold indicates that the speaker's speaking time has not yet reached the set duration limit.
[0037] It should be noted that the first and second thresholds are thresholds set based on application requirements, and there are no restrictions on their values.
[0038] The first cue message can be provided only to the current speaker, or it can be provided to the meeting facilitator, who will then determine whether to provide it to the current speaker. The first cue message can take any form and is used to help speakers understand which key information has not yet been mentioned, so that they can supplement it in their subsequent remarks.
[0039] In one example, in a project reporting meeting, the host or organizer pre-sets the required content structure for the session, including the following key points: pain points, solutions, implementation paths, risks, and a time limit for speaking. The time limit can be set using any of the first, second, or third restriction methods mentioned above, which will not be elaborated here.
[0040] During the conversation, a speaker is delivering a speech. Conversation data is collected in real time, including the content and duration of the speaker's speech. This data is then parsed, and the parsing results are assumed to include: The content of the speech did not reach the first threshold. For example, the speech covered the pain points and solutions, but did not mention the implementation path and risks; or the speech covered the pain points and implementation path, but omitted the explanation of the solutions. The speaking time did not reach the second threshold. For example, the speaking time is 20 minutes, but the meeting stipulates that the maximum time for each speaker is 30 minutes.
[0041] At this point, a first prompt message can be generated. This first prompt message is used to guide the current speaker to speak based on the content in the first prompt message to improve the coverage of the speech content. For example, the first prompt message may include: Please explain the implementation path, risks, etc. This content is provided to the current speaker. After receiving the first prompt message, the speaker understands and needs to supplement the missing information, and then explains the missing content in detail in the following speech, thereby improving the coverage of the speech content.
[0042] It should be noted that the first prompt message can be sent only to the speaker, or it can be sent to the moderator, with the moderator deciding whether to send it to the speaker. Alternatively, it can be sent to both the speaker and the moderator simultaneously, so that the moderator can be informed of the meeting's progress in a timely manner and to facilitate verbal reminders from the moderator.
[0043] In this way, the system can guide speakers to supplement important information, improve the overall efficiency of the meeting, and ensure that each speaker's remarks are more comprehensive and effective.
[0044] In some embodiments, the processing logic is determined based on the parsing result of the session data, including: In response to the session data satisfying a third condition, a third processing logic is determined, the third processing logic including generating a second prompt message based on the session data and outputting the second prompt message to the current speaker; The third condition indicates that the content coverage of the conversation data does not reach the first threshold, the speaking duration does not reach the second threshold, and the remaining speaking time predicted based on the conversation data exceeds the third threshold. The second prompt message is used to guide the current speaker to speak based on the content of the second prompt message in order to improve the coverage of the speech content, and to indicate the remaining speaking time.
[0045] Here, the third condition includes two aspects: the coverage of the speech content and the duration of the speech.
[0046] The coverage of the speech content refers to the comprehensiveness and sufficiency of the key points or information involved in the speech; the coverage of the speech content not reaching the first threshold indicates that the current speaker has not fully covered all the necessary topics or information points in the speech.
[0047] The speaking duration represents the time currently used by the speaker. The speaker not reaching the second threshold indicates that the speaker's speaking time has not yet reached the set duration limit.
[0048] The predicted remaining speaking time represents the time the system predicts the speaker will need to fully cover the necessary information. If the predicted remaining speaking time exceeds a third threshold, it means the speaker needs a longer time to complete the speech, exceeding a set threshold.
[0049] It should be noted that the first threshold, the second threshold, and the third threshold are thresholds set based on application requirements, and there are no restrictions on their values.
[0050] The second cue message can be provided only to the current speaker, or it can be provided to the meeting moderator, who will then decide whether to provide it to the current speaker. The second cue message can guide the speaker to supplement any missing key information, and can also provide feedback on the remaining speaking time so that the speaker can plan their subsequent remarks accordingly.
[0051] In one example, in a project reporting meeting, the host or organizer pre-sets the required content structure for the session, including the following key points: pain points, solutions, implementation paths, risks, and a time limit for speaking. The time limit can be set using any of the first, second, or third restriction methods mentioned above, which will not be elaborated here.
[0052] During the conversation, a speaker is delivering a speech. Conversation data is collected in real time, including the content and duration of the speaker's speech. This data is then parsed, and the parsing results are assumed to include: The content of the speech did not reach the first threshold. For example, the speech covered the pain points and solutions, but did not mention the implementation path and risks; or the speech covered the pain points and implementation path, but omitted the explanation of the solutions. The speaking time did not reach the second threshold. For example, the speaking time is 20 minutes, but the meeting stipulates that the maximum time for each speaker is 30 minutes. If the predicted remaining speaking time exceeds the third threshold, for example, the system predicts that the speaker needs at least 20 minutes of speaking time to cover all the necessary information, which exceeds the third threshold. The third threshold can be the reasonable meeting duration minus the currently used time (e.g., 10 minutes obtained by subtracting 30 minutes and 20 minutes as described above). Another example is that the speaking content has covered pain points and solutions, but has not yet mentioned implementation paths and risks. The system can predict separately for each uncovered point, such as the speaker needing at least 10 minutes to explain the implementation path and 5 minutes to explain the risks.
[0053] At this point, if the third condition is met, the system can generate a second prompt message, such as: "Please provide additional information about the implementation path and risks. You have 10 minutes left to speak."
[0054] The second notification message can be sent only to the speaker, or it can be sent to both the speaker and the moderator at the same time, so that the moderator can be informed of the meeting situation in a timely manner, or it can facilitate the moderator to give verbal reminders, etc.
[0055] After receiving the second prompt, the speaker can understand the information that needs to be supplemented and use the remaining speaking time to quickly explain it in the following speech, so as to reduce the required time while providing a complete explanation.
[0056] After receiving the second notification, the host can promptly understand that there may be errors in the time allocation and make corresponding offline adjustments based on the actual situation; no restrictions are imposed here.
[0057] In this way, not only were speakers guided to increase the coverage of their speeches, but the structure of their speeches was also ensured to be reasonable, making the discussions at the meeting more comprehensive and efficient.
[0058] In some embodiments, parsing the session data based on the model includes: The model identifies the content of the conversation data and analyzes the key points involved in the content. The content coverage is determined based on the analysis results. The content coverage represents the degree to which the key points involved in the content cover the target key points. Determine the speaking duration of the session data.
[0059] Here, the model can be used to identify the content of a speech and analyze it to extract key points. Specific techniques that can be used include, but are not limited to, semantic recognition, keyword extraction, and key point analysis. Once the key points are identified, the relationship between the extracted key points and predefined target key points (i.e., the expected content of the meeting discussion) can be compared to calculate the coverage of the speech content. This coverage reflects whether the speaker's argument is comprehensive and whether it covers all the necessary information.
[0060] The coverage of the speech content can be expressed as a percentage. For example, if the target points include pain points, solutions, implementation paths, and risks, then the coverage is 50% if the speaker only covers 50% of the target points (such as pain points and solutions).
[0061] Speaking duration refers to the actual length of time spoken during the meeting, which can be determined by analyzing audio files or text transcripts.
[0062] In one example, during a conversation, a speaker is speaking, and conversation data is collected in real time. The collected conversation data includes audio recordings, which can be semantically analyzed to determine the content of the speech.
[0063] For example, if the speech content states, "Currently, there is the following problem 1 regarding xx technology (specific examples are not given here), and a solution 1 is proposed to address this problem (specific examples are not given here)," and the analysis confirms that the pain point and solution have been explained, but the implementation path and risks have not been addressed, then the current coverage is determined to be 50%. As another example, if there are 5 key objectives, and the speech content analysis determines that only 2 key objectives are covered, then the coverage is 40%. The speech duration can be calculated using audio recordings, such as the difference between the start time and the current time.
[0064] In this way, by effectively identifying and analyzing the content of speeches through the model, calculating the coverage and duration of speeches, and gaining a comprehensive understanding of the progress of the conversation, the conversation process can be improved through various processing logics, guiding the speakers' statements and ensuring the comprehensiveness and effectiveness of the meeting discussion.
[0065] In some embodiments, if the coverage of the spoken content exceeds a third threshold, and the third threshold is less than the first threshold, the method further includes: In response to the session data satisfying the fourth condition, determine the second or third processing logic; The session data satisfies a fourth condition, including at least one of the following: The correlation between any two adjacent key points in the content of the conversation data does not exceed the fourth threshold. The relevance between the content of the conversation data and the real-time shared file corresponding to the current conversation does not exceed the fifth threshold.
[0066] Here, considering that even if a speaker's speech covers a certain extent, the key points may not be clearly described. For example, although the "pain point" is mentioned, the described "pain point" has a low correlation with the subsequent "solution," making it difficult for others to understand the actual content and the interrelationships. Therefore, in the process of managing the conversation, this embodiment of the disclosure can further determine whether the speech content is adequate and whether each key point is clearly described.
[0067] The third threshold can be a lower standard used to trigger a judgment on whether the session data meets the fourth condition. Specifically, considering that the judgment on whether the session data meets the fourth condition requires a certain amount of speech content, it is proposed to make a judgment when the speech content coverage exceeds the third threshold.
[0068] The third, fourth, and fifth thresholds can be set based on actual needs, and there are no restrictions on their values.
[0069] If the session data meets the fourth condition, the second processing logic can be executed to output the first prompt message to the speaker, or the third processing logic can be executed to output the second prompt message to the speaker. The first and second prompt messages can indicate that a certain point is not clearly described, or that the point does not match the shared file, so as to guide the current speaker to make adjustments.
[0070] The choice between executing the second or third processing logic can be made based on whether the session data meets the second or third condition. That is, if the session data meets both the second and fourth conditions, the second processing logic is executed; if the session data meets both the third and fourth conditions, the third processing logic is executed.
[0071] The determination of whether the fourth condition is met can include at least one of the following: Based on the key points involved in the speech content of the conversation data, a key point correlation analysis is performed to determine whether the correlation between any two adjacent key points exceeds the fourth threshold. If it does not exceed the fourth threshold, then the fourth condition is satisfied. The relevance analysis is performed between the spoken content of the conversation data and the real-time shared file corresponding to the current speech. If the relevance does not exceed the fifth threshold, the fourth condition is determined to be met. The real-time shared file corresponding to the current speech can be a descriptive file about the meeting content, such as a PPT, Word document, video, or image.
[0072] If the correlation between adjacent key points does not exceed the fourth threshold, it is considered that the logical or thematic connection between any two adjacent key points in the speech is weak, indicating that the speech lacks coherence, and the description is considered unclear.
[0073] If the relevance of the speaker's content to the real-time shared files does not exceed the fifth threshold, it is considered that the content being discussed by the speaker is not sufficiently relevant to the presented materials or files, which may lead to inconsistencies in information delivery or unclear descriptions.
[0074] In one example, suppose an engineer is reporting on project progress during a technical discussion meeting. After collecting a certain amount of conversation data and performing a content coverage analysis, the content coverage is determined to be 40%, reaching the third threshold of 40% (assuming the first threshold is 50%). Further judgment is then made as to whether the conversation data meets the fourth condition.
[0075] For example, the key points obtained from the analysis include: Key point 1: "Problem 1"; Key point 2: "Solution 1 proposed for Problem 1"; Key point 3: "Implementation path 1 of Solution 1". After analysis, the correlation between key point 1 and key point 2 is high, and solution 1 is specifically proposed for problem 1. The correlation between key point 3 and key point 2 is also high, and implementation path 1 is specifically proposed for solution 1. Therefore, it can be considered that the fourth condition is met.
[0076] Conversely, if the key points obtained from the analysis include: Key Point 1: "Problem 1"; Key Point 2: "Solution 2 proposed for Problem 1"; Key Point 3: "Implementation path 3 of Solution 2", and after analysis, it is found that Solution 2 is not specifically proposed for Problem 1, the correlation between Key Point 1 and Key Point 2 is low, and / or Key Point 3 is also not highly correlated with Key Point 2, and Implementation path 3 is not specifically proposed for Solution 2, then it can be considered that the fourth condition is not met.
[0077] For example, analyzing the relevance between the speech content and the shared file can be done by comparing the speech content with a document shown in the meeting about the current project development progress. If the comparison finds that the actual content of the two does not match, such as the speech content being Question 1, while the content presented in the document is Question 2 or something else, then the relevance between the speech content and the real-time shared file is considered to be no more than the fifth threshold.
[0078] In this case, the second processing logic is executed to output the first prompt message to the speaker, or the third processing logic is executed to output the second prompt message to the speaker. The first and second prompt messages may indicate that a certain point is not clearly described, or that the point does not match the shared file, so as to guide the current speaker to make adjustments.
[0079] In this way, not only can the coverage of the speech be assessed, but adjustments can also be made based on the coherence and relevance of the content, thereby improving the quality and effectiveness of the speech and helping to ensure that the information discussed at the meeting is more comprehensive and in-depth.
[0080] In some embodiments, generating a first prompt message based on the session data includes: Determine the key points involved in the spoken content of the conversation data; The key points mentioned in the speech are compared with the target key points to identify the target key points that are not covered. A first prompt message is generated based on the uncovered target points.
[0081] Here, the content of the first prompt can be determined by analyzing the speech content in the conversation data, so as to more comprehensively cover the relevant content in subsequent meetings.
[0082] Specifically, by analyzing conversation data (e.g., speech-to-text), the key points mentioned by the speaker in their speech are identified, and the key points involved in the speech content are compared with predefined target key points to determine the uncovered target key points. The first prompt information is used to indicate the uncovered target key points.
[0083] For example, the speech may cover the pain points and solutions, but not the implementation path and risks; or the speech may cover the pain points and implementation path, but omit the explanation of the solutions. In such cases, a first prompt message can be generated, including: Please explain the implementation path and risks; or, Please supplement the solution content. This content is provided to the current speaker. After receiving the first prompt message, the speaker understands and needs to supplement the missing information, and then explains the missing content in detail in the following speech, thereby improving the coverage of the speech content.
[0084] Similarly, the second prompt message can be determined based on the uncovered target points and the predicted remaining speaking time. For example, the speaking content has covered pain points and solutions, but has not yet mentioned implementation paths and risks; the speaking time is 20 minutes, but the meeting stipulates that the maximum time for each speaker is 30 minutes, and the predicted remaining speaking time exceeds the remaining 10 minutes, the generated second prompt message could include: "Please add information about implementation paths and risks. You have 10 minutes left to speak, please be mindful of the time limit."
[0085] In this way, by identifying and comparing the content of speeches with the target points, speakers are helped to realize important information that may have been omitted in their presentations, thereby improving the comprehensiveness and effectiveness of future discussions. Time reminders can also help speakers manage their time effectively, thus improving meeting efficiency.
[0086] In some embodiments, predicting the remaining duration of a speech based on the session data includes: The conversation rhythm is determined based on the conversation data, and the conversation rhythm includes the current conversation progress and / or whether the conversation is currently in a discussion state. Based on the conversation data, a feature vector of the current speaker is extracted, and the feature vector includes at least one of the following features: speech rate, pause rhythm, skipping habits, and content depth; The remaining duration required for the speech is determined based on the conversation rhythm, the feature vector, and the duration of the remaining speech content.
[0087] Here, the pace of the conversation includes the speaker's progress and whether they are in a discussion. The current conversation progress can include the key points described and the time elapsed; whether they are in a discussion indicates whether there is a multi-person discussion in the meeting, which takes into account that multi-person discussions may slow down the pace, while individual speeches may be faster.
[0088] Speech rate refers to how many words a speaker speaks per minute, which directly affects the duration of a speech.
[0089] Pause rhythm can refer to the frequency and duration of a speaker's pauses during a speech. More pauses may lead to a longer overall speaking time.
[0090] Skimming refers to a speaker's tendency to quickly browse or skip certain content, which can affect the completeness and duration of their speech.
[0091] Content depth refers to the complexity and depth of the speech content; more complex content usually requires more time to explain and discuss.
[0092] By combining the conversation rhythm, the speaker's feature vector, and the duration of the remaining speech, the time required for the speaker to complete the remaining speech can be calculated.
[0093] In one example, the remaining speaking time TA can be predicted using the following formula: Remaining time (TA) = Remaining content volume × Speaking speed factor × Meeting pace factor Among them, the remaining content represents the duration of the key points that have not yet been finished. This duration can be the theoretical speaking time set by the host for each key point when setting the key points. The speaking speed factor can be based on dynamic behaviors such as the speaker's speaking speed, pause rhythm, skipping habits, and content depth. For example, it can be calculated using the following formula: Speaking speed factor (SpeedFactor) = (Standard speaking speed / Current speaker's speaking speed) (1 + pause ratio) (1 + percentage of questions skipped) (1+ content depth); The conversation pace factor can be represented by PaceFactor, which is determined based on certain rules. For example, if 70% of the time has already been spent, then PaceFactor = 70%; if the discussion phase has begun, then PaceFactor = 1. Discuss weights, such as PaceFactor=1.2.
[0094] In another example, the conversation rhythm, speaker's feature vector, and the duration of remaining speech content can be input into the prediction model to obtain the required remaining duration TB output by the prediction model. The required remaining duration TA is then adjusted based on the required remaining duration TB output by the prediction model and weights to obtain the final required remaining duration. The weights represent the confidence level of the remaining duration TA. The prediction model can be trained based on historical data to dynamically adjust the prediction results. For example, the final required remaining duration T... final =λ TA+ (1-λ) TB; where TA represents the predicted remaining time; TB represents the predicted remaining time obtained by the prediction model; λ∈[0,1] represents the weight, which is adjustable, and the more reliable the prediction, the larger λ is.
[0095] Of course, other methods can be used to predict the remaining time required, which are not limited here.
[0096] In some embodiments, determining the processing logic based on the parsing result of the session data further includes: In response to the session data satisfying the second, third, or fifth condition, a fourth processing logic is determined, the fourth processing logic including determining a third prompt message based on the real-time shared file corresponding to the current speech, and outputting the third prompt message to other participants in the session; The second condition indicates that the content coverage of the conversation data does not reach the first threshold and the speaking duration does not reach the second threshold. The third condition indicates that the content coverage of the conversation data does not reach the first threshold, the speaking duration does not reach the second threshold, and the remaining speaking time predicted based on the conversation data exceeds the third threshold. The fifth condition indicates that the coverage of the speech content in the conversation data has not reached the first threshold and the speech duration has reached the second threshold. The third prompt is used to help other participants in the session understand the content of the current speaker's speech.
[0097] Regarding the second and third conditions, please refer to the above explanation; they will not be repeated here. As for the fifth condition, if the content coverage does not reach the first threshold but the speaking time reaches the second threshold, it indicates that the speaker spent a considerable amount of time, but the content is still insufficient to cover the key points or is inadequately described. All of the above situations may prevent other participants from understanding the relevant project information. Therefore, to facilitate a better understanding of the speaking content for other participants, a third prompt can be used for clarification.
[0098] The purpose of the third cue is to help other participants in the meeting better understand the current speaker's content. This can include summaries of key points, supplementary information, explanatory information, etc., to ensure that participants can keep up with the progress of the conversation.
[0099] In one example, the speaker outlined the key point "Option 1," but the steps within Option 1 were not fully described. The third prompt could be a flowchart or process description of Option 1 generated based on analysis of the shared file in real-time. In another example, the project manager is presenting the progress of a new project, but the content is incomplete. The third prompt could be detailed progress information determined from the file; if this is not explicitly marked in the file, it could suggest, "Information regarding project progress is missing; please refer to subsequent discussions for further information." In yet another example, a technical expert used numerous technical terms when explaining a new technology; the third prompt could be an explanation of these technical terms.
[0100] In this way, by analyzing the conversation data, third-party prompts can be generated for the participants in specific situations to help enhance their understanding of the content, thereby improving the effectiveness of the meeting, ensuring that all participants receive the necessary information during the discussion, and deepening their understanding of the content.
[0101] Figure 2 A flowchart illustrating a meeting method based on content coverage and time prediction, provided for an application embodiment of this disclosure; as shown Figure 2As shown, the method may include: at the start of a project meeting, the facilitator may pre-set the meeting structure information, which may include: the target points involved in each round of the conversation (such as pain points, solutions, implementation paths, risks, and next steps), and the set duration of each round of the conversation (which may be set for the entire conversation or for each target point individually); the meeting structure information may also include the setting of a first threshold, a second threshold, a third threshold, a fourth threshold, and a fifth threshold.
[0102] For each round of conversation (e.g., a conversation of speaker 1, speaker 2... speaker N), after a speaker begins speaking, conversation data is collected. The coverage and duration of the speech content are analyzed based on the conversation data, and corresponding processing logic is executed based on the analysis results until the speaker finishes speaking and enters the next round. This processing logic may include a first processing logic, a second processing logic, and a third processing logic, specifically as follows: Figure 1 The explanations for the methods shown will not be repeated here.
[0103] Furthermore, for each speaker's remarks, a corresponding structured summary of the project and a structured summary of the entire meeting can be generated to collect and organize the remarks for quick review later.
[0104] Figure 3 A flowchart illustrating a session data processing method provided in an embodiment of this disclosure; as shown below. Figure 3 As shown, the method includes: Step 301: Collect session data; Here, the session data includes at least the content and duration of the current speaker's speech.
[0105] Step 302: Parse the session data; Here, the model can be used to parse the conversation data, which may specifically include: identifying the content of the conversation data, analyzing the key points involved in the content of the conversation, and determining the duration of the conversation data.
[0106] Step 303: Analyze the analysis results; Here, based on the parsing results, it can be determined whether the content and duration of the speech match the set conditions. Based on the matching results, corresponding processing logic is executed, which may include: If the content of the speech is insufficient, a prompt message will be generated to supplement the key points; If time is insufficient, a prompt message will be generated to accelerate or compress the expression; If the content and duration of the speech meet the requirements, monitoring will continue.
[0107] In this case, insufficient content in the speech can refer to the fact that the key points involved in the speech do not fully cover all the target key points. Accordingly, relevant prompts for insufficient content in the speech will be given, such as generating prompts to supplement key content (such as target key points that are not covered). Insufficient time can refer to the predicted remaining time for the speech exceeding a certain threshold. In response, relevant prompts regarding insufficient time will be given, such as generating prompts to accelerate or compress the expression.
[0108] Step 304: Output the determined processing logic as a visual result; Here, the visualization results can be output to the conversation assistant for decision-making, such as sending the above prompts to the speaker; of course, the visualization results can also be sent directly to the speaker.
[0109] Figure 4 A flowchart illustrating a session data processing method provided in an embodiment of this disclosure; as shown below. Figure 4 As shown, the method includes: Step 401: Determine the key points that match the content of the speech; Step 402: Determine the content coverage level; Here, the degree of content coverage represents the extent to which the key points involved in the speech cover the target key points. For example, whether the key points are completely hit, whether the key points are partially hit, or whether the key points are not hit at all. Assuming the target points include: Target Point 1, Target Point 2, and Target Point 3, if the speech content involves only one or two points, such as Target Point 1 and Target Point 2 matching Target Point 1 and Target Point 2, but Target Point 3 is not covered, it is considered a partial hit; if all target points are successfully matched, it is considered a complete hit; if none are successfully matched, it is considered a miss.
[0110] Step 403: Summarize the key points covered and provide corresponding prompts; Here, by summarizing the covered key points, we can identify the uncovered key points. For the uncovered key points, we can provide content prompts to guide the current speaker to speak based on the prompts, thereby improving the coverage of the speech content.
[0111] Step 404: Predict the remaining speaking time based on the remaining content and provide corresponding prompts; Here, the remaining content is determined based on the uncovered key points. The remaining content duration, the speaker's feature vector, and the meeting pace are then combined to predict the remaining speaking time. The meeting pace includes the current conversation progress and / or whether the discussion is ongoing. The feature vector includes at least one of the following features: speaking speed, pause rhythm, skipping habits, and content depth. If the predicted remaining speaking time exceeds a set threshold, a prompt is given to indicate the remaining speaking time and guide the speaker to accelerate or compress their expression.
[0112] Figure 5 This is a schematic diagram of the structure of a session data processing apparatus according to an embodiment of the present disclosure; as shown below. Figure 5 As shown, the device includes: The acquisition module is used to acquire session data during the session; the session data includes at least the content and duration of the current speaker's speech. The processing module is used to parse the session data based on the model; determine the processing logic based on the parsing result of the session data; and execute the processing logic to control the progress of the session.
[0113] In some embodiments, the processing module is configured to determine first processing logic in response to the session data satisfying a first condition, wherein the first processing logic is configured to terminate the current speech; The first condition indicates that the content coverage of the conversation data reaches a first threshold and the speaking duration does not reach a second threshold.
[0114] In some embodiments, the processing module is configured to determine second processing logic in response to the session data satisfying a second condition, the second processing logic including generating a first prompt message based on the session data and outputting the first prompt message to the current speaker; The second condition indicates that the content coverage of the conversation data does not reach the first threshold and the speaking duration does not reach the second threshold. The first prompt message is used to guide the current speaker to speak based on the content of the first prompt message in order to improve the coverage of the speech content.
[0115] In some embodiments, the processing module is configured to determine third processing logic in response to the session data satisfying a third condition, the third processing logic including generating a second prompt message based on the session data and outputting the second prompt message to the current speaker; The third condition indicates that the content coverage of the conversation data does not reach the first threshold, the speaking duration does not reach the second threshold, and the remaining speaking time predicted based on the conversation data exceeds the third threshold. The second prompt message is used to guide the current speaker to speak based on the content of the second prompt message in order to improve the coverage of the speech content, and to indicate the remaining speaking time.
[0116] In some embodiments, the processing module is configured to identify the speech content of the session data based on a model, analyze the key points involved in the speech content, and determine the speech content coverage based on the analysis results; the speech content coverage represents the degree to which the key points involved in the speech content cover the target key points; Determine the speaking duration of the session data.
[0117] In some embodiments, if the coverage of the spoken content exceeds a third threshold, and the third threshold is less than the first threshold, the processing module is configured to determine a second processing logic or a third processing logic in response to the session data satisfying a fourth condition. The session data satisfies a fourth condition, including at least one of the following: The correlation between any two adjacent key points in the content of the conversation data does not exceed the fourth threshold. The relevance between the content of the conversation data and the real-time shared file corresponding to the current conversation does not exceed the fifth threshold.
[0118] In some embodiments, the processing module is configured to determine the key points involved in the spoken content of the session data; The key points mentioned in the speech are compared with the target key points to identify the target key points that are not covered. A first prompt message is generated based on the uncovered target points.
[0119] In some embodiments, the processing module is configured to determine the session rhythm based on the session data, wherein the session rhythm includes the current session progress and / or whether the current state is in discussion. Based on the conversation data, a feature vector of the current speaker is extracted, and the feature vector includes at least one of the following features: speech rate, pause rhythm, skipping habits, and content depth; The remaining duration required for the speech is determined based on the conversation rhythm, the feature vector, and the duration of the remaining speech content.
[0120] In some embodiments, the processing module is further configured to determine a fourth processing logic in response to the session data satisfying a second condition, a third condition, or a fifth condition. The fourth processing logic includes determining a third prompt message based on the real-time shared file corresponding to the current speech and outputting the third prompt message to other participants in the session. The second condition indicates that the content coverage of the conversation data does not reach the first threshold and the speaking duration does not reach the second threshold. The third condition indicates that the content coverage of the conversation data does not reach the first threshold, the speaking duration does not reach the second threshold, and the remaining speaking time predicted based on the conversation data exceeds the third threshold. The fifth condition indicates that the coverage of the speech content in the conversation data has not reached the first threshold and the speech duration has reached the second threshold. The third prompt is used to help other participants in the session understand the content of the current speaker's speech.
[0121] It is understood that, when implementing the corresponding session data processing method, the session data processing apparatus provided in the above embodiments can allocate the processing to different program modules as needed to complete all or part of the processing described above. Furthermore, the apparatus and the corresponding method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0122] This disclosure provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored and, when executed by a processor, will trigger the processor to execute the session data processing method provided in this disclosure.
[0123] In some embodiments, the computer-readable storage medium may be a ferroelectric random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disc, or CD-ROM, etc.; or it may be a device that includes one or any combination of the above-mentioned memories.
[0124] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, model, subroutine, or other unit suitable for use in a computing environment.
[0125] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0126] This disclosure provides a computer program product, which includes a computer program / instruction that, when executed by a processor, implements the session data processing method described in this disclosure.
[0127] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure; as shown below. Figure 6 As shown, the electronic device 60 includes: a processor 601 and a memory 602 for storing computer programs that can run on the processor; when the processor 601 runs the computer program, it executes the session data processing method provided in the embodiments of this disclosure.
[0128] In practical applications, the electronic device 60 may further include at least one network interface 603. The various components of the electronic device 60 are coupled together via a bus system 604. It is understood that the bus system 604 is used to implement communication between these components. In addition to a data bus, the bus system 604 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 6 Various buses are designated as bus system 604. The number of processors 601 can be at least one. Network interface 603 is used for wired or wireless communication between electronic device 60 and other devices.
[0129] The memory 602 in this embodiment is used to store various types of data to support the operation of the electronic device 60.
[0130] The methods disclosed in the above embodiments of this disclosure can be applied to processor 601, or implemented by processor 601. Processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 601 or by instructions in the form of software. The processor 601 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 601 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory 602. Processor 601 reads the information in memory 602 and combines its hardware to complete the steps of the aforementioned method.
[0131] In some embodiments, the electronic device 60 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned methods.
[0132] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0133] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0134] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A session data processing method, the method comprising: During the session, obtain session data; The session data includes at least the content and duration of the current speaker's speech; The session data is parsed based on the model; The processing logic is determined based on the parsing results of the session data, and the processing logic is executed to control the progress of the session.
2. The method according to claim 1, wherein the processing logic is determined based on the parsing result of the session data, comprising: In response to the session data satisfying a first condition, a first processing logic is determined, which is used to end the current speech; The first condition indicates that the content coverage of the conversation data reaches a first threshold and the speaking duration does not reach a second threshold.
3. The method according to claim 1, wherein the processing logic is determined based on the parsing result of the session data, comprising: In response to the session data satisfying the second condition, a second processing logic is determined, the second processing logic including generating a first prompt message based on the session data and outputting the first prompt message to the current speaker; The second condition indicates that the content coverage of the conversation data does not reach the first threshold and the speaking duration does not reach the second threshold. The first prompt message is used to guide the current speaker to speak based on the content of the first prompt message in order to improve the coverage of the speech content.
4. The method according to claim 1, wherein the processing logic is determined based on the parsing result of the session data, comprising: In response to the session data satisfying a third condition, a third processing logic is determined, the third processing logic including generating a second prompt message based on the session data and outputting the second prompt message to the current speaker; The third condition indicates that the content coverage of the conversation data does not reach the first threshold, the speaking duration does not reach the second threshold, and the remaining speaking time predicted based on the conversation data exceeds the third threshold. The second prompt message is used to guide the current speaker to speak based on the content of the second prompt message in order to improve the coverage of the speech content, and to indicate the remaining speaking time.
5. The method according to claim 1, wherein parsing the session data based on the model includes: The model identifies the content of the conversation data and analyzes the key points involved in the content of the conversation. The coverage of the speech content is determined based on the analysis results; The coverage of the speech content represents the degree to which the key points involved in the speech content cover the target key points; Determine the speaking duration of the session data.
6. The method according to claim 3 or 4, wherein if the coverage of the spoken content exceeds a third threshold, and the third threshold is less than the first threshold, the method further comprises: In response to the session data satisfying the fourth condition, determine the second or third processing logic; The session data satisfies a fourth condition, including at least one of the following: The correlation between any two adjacent key points in the content of the conversation data does not exceed the fourth threshold. The relevance between the content of the conversation data and the real-time shared file corresponding to the current conversation does not exceed the fifth threshold.
7. The method according to claim 3, wherein generating a first prompt message based on the session data includes: Determine the key points involved in the spoken content of the conversation data; The key points mentioned in the speech are compared with the target key points to identify the target key points that are not covered. A first prompt message is generated based on the uncovered target points.
8. The method according to claim 4, wherein predicting the remaining duration of the speech based on the session data includes: The conversation rhythm is determined based on the conversation data, and the conversation rhythm includes the current conversation progress and / or whether the conversation is currently in a discussion state. Based on the conversation data, a feature vector of the current speaker is extracted, and the feature vector includes at least one of the following features: speech rate, pause rhythm, skipping habits, and content depth; The remaining duration required for the speech is determined based on the conversation rhythm, the feature vector, and the duration of the remaining speech content.
9. The method according to claim 1, further comprising determining processing logic based on the parsing result of the session data: In response to the session data satisfying the second, third, or fifth condition, a fourth processing logic is determined, the fourth processing logic including determining a third prompt message based on the real-time shared file corresponding to the current speech, and outputting the third prompt message to other participants in the session; The second condition indicates that the content coverage of the conversation data does not reach the first threshold and the speaking duration does not reach the second threshold. The third condition indicates that the content coverage of the conversation data does not reach the first threshold, the speaking duration does not reach the second threshold, and the remaining speaking time predicted based on the conversation data exceeds the third threshold. The fifth condition indicates that the coverage of the speech content in the conversation data has not reached the first threshold and the speech duration has reached the second threshold. The third prompt is used to help other participants in the session understand the content of the current speaker's speech.
10. An electronic device, comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform: During the conversation, conversation data is obtained; the conversation data includes at least the content and duration of the current speaker's speech. The session data is parsed based on the model; processing logic is determined based on the parsing results of the session data, and the processing logic is executed to control the progress of the session.