Time analysis in video conference

A system using a large language model to allocate session durations and visual/audio cues in video conferencing systems addresses inefficiencies by managing session times effectively, improving conference organization and user experience.

WO2025254722A1PCT designated stage Publication Date: 2025-12-11MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/022643
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-05
Filing Date
2025-04-02
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing video conferencing systems lack effective methods for automatically allocating and managing session durations, leading to inefficiencies and potential interruptions during conferences.

Method used

Implement a system that uses a large language model to analyze conference descriptions and documents to determine session durations, and employs timers with audio and visual cues to manage session times, allowing speakers to adjust their speaking time.

Benefits of technology

This approach enables efficient session duration management, reducing the need for human intervention and minimizing abrupt interruptions while enhancing conference organization and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025022643_11122025_PF_FP_ABST
    Figure US2025022643_11122025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure proposes a method, apparatus and computer program product for time analysis in video conference. A total duration of a video conference may be obtained. A conference description of the video conference and / or a document associated with the video conference may be obtained. A plurality of sessions included in the video conference and a planned duration of each session may be determined based on at least one of the total duration, the conference description, and the document. For each session, it may be detected that the video conference proceeds to the session. In response to detecting that the video conference proceeds to the session, a timer corresponding to the session may be started, the timer displaying a remaining duration associated with a planned duration of the session. In response to the remaining duration being below a predetermined threshold, an indication may be output.
Need to check novelty before this filing date? Find Prior Art

Description

TIME ANALYSIS IN VIDEO CONFERENCEBACKGROUND

[0001] With the development of digital device, communication technology, video processing technology, etc., people may use terminal devices such as desktop computer, tablet computer, smart phone, etc., to conduct video conference with people located elsewhere for purposes such as work discussions, remote training, technical support, etc. Herein, video conference may broadly refer to a conference method based on Internet technology that can transmit voice and images of participants in real time. Video conference may also be referred to as online conference, virtual conference, etc.SUMMARY

[0002] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. It is not intended to identity key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0003] Embodiments of the present disclosure propose a method, apparatus and computer program product for time analysis in video conference. A total duration of a video conference may be obtained. A conference description of the video conference and / or a document associated with the video conference may be obtained. A plurality of sessions included in the video conference and a planned duration of each session may be determined based on at least one of the total duration, the conference description, and the document. For each session in the plurality of sessions, it may be detected that the video conference proceeds to the session. In response to detecting that the video conference proceeds to the session, a timer corresponding to the session may be started, the timer displaying a remaining duration associated with a planned duration of the session. In response to the remaining duration being below a predetermined threshold, an indication may be output, the indication including an audio signal and / or a visual effect.

[0004] It should be noted that the above one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the drawings set forth in detail certain illustrative features of the one or more aspects. These features are only indicative of the various ways in which the principles of various aspects may be employed, and this disclosure is intended to include all such aspects and their equivalents.BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The disclosed aspects will hereinafter be described in connection with the appended drawings that are provided to illustrate and not to limit the disclosed aspects.

[0006] FIG. 1 illustrates an exemplary process for time analysis in video conference accordingto an embodiment of the present disclosure.

[0007] FIG. 2 illustrates an exemplary process for determining a set of content sessions included in a video conference and a planned duration of each content session with a conference description of the video conference according to an embodiment of the present disclosure.

[0008] FIG. 3 illustrates an exemplary process for determining a set of content sessions included in a video conference and a planned duration of each content session with a document associated with the video conference according to an embodiment of the present disclosure.

[0009] FIG. 4 illustrates an example of outputting a visual effect according to an embodiment of the present disclosure.

[0010] FIG. 5 illustrates another example of outputting a visual effect according to an embodiment of the present disclosure.

[0011] FIG. 6 is a flowchart of an exemplary method for time analysis in video conference according to an embodiment of the present disclosure.

[0012] FIG. 7 illustrates an exemplary apparatus for time analysis in video conference according to an embodiment of the present disclosure.

[0013] FIG. 8 illustrates another exemplary' apparatus for time analysis in video conference according to an embodiment of the present disclosure.DETAILED DESCRIPTION

[0014] The present disclosure will now be discussed with reference to several example implementations. It is to be understood that these implementations are discussed only for enabling those skilled in the art to better understand and thus implement the embodiments of the present disclosure, rather than suggesting any limitations on the scope of the present disclosure.

[0015] A video conference usually includes a plurality of sessions, such as an opening session for generally introducing topics to be discussed in the video conference, goals to be achieved in the video conference, etc., a set of content sessions for elaborating or discussing the conference topics, an ending session for questions and answers, summarization, etc. When there are a plurality’ of topics that need to be elaborated or discussed, there will be a plurality of content sessions accordingly. Different content sessions may be elaborated or responsible by different participants of the conference. It is desirable to reasonably allocate the time of each session to complete the elaboration or discussion of all topics within the predetermined conference time.

[0016] Embodiments of the present disclosure propose to perform time analysis on a video conference, to automatically determine a plurality of sessions of the video conference and a planned duration of each session. A planned duration of an opening session may be automatically calculated based on a predetermined total duration of the video conference and a first predetermined ratio. A planned duration of an ending session may be automatically calculatedbased on the total duration and a second predetermined ratio. The second predetermined ratio may be the same as or different from the first predetermined ratio. A set of content sessions for elaborating or discussing conference topics and a planned duration of each content session may be automatically determined in a variety of ways. In an implementation, the set of content sessions of the video conference and the planned duration of each content session may be automatically determined based on the total duration and a conference description of the video conference through a Large Language Model (LLM). Herein, a language model refers to a deep learning model that can understand meaning of natural language, generate natural language texts, or perform other natural language tasks. It should be appreciated that the large language models include multi-modal models that can perform processing tasks for natural language as well as other modalities. The conference description may be extracted from a conference invitation email of the video conference, which may include, e.g., a set of topics planned to be discussed in the video conference, a set of speakers to speak during the video conference, etc. In another implementation, the set of content sessions included in the video conference and the planned duration of each content session may be automatically determined with a document associated with the video conference. The document associated with the video conference may be a document attached to the conference invitation email of the video conference, a document shared during the video conference, etc. The document includes a plurality of pages, and a duration of each page may be determined based on content of the page. The set of content sessions may be determined through analyzing a set of topics included in the document. For each content session, a set of pages belonging to the content session may be identified, and a planned duration of the content session may be calculated based on a set of durations corresponding to the set of pages. The technical effects of the implementation described above are that the duration of each session of the video conference can be reasonably allocated, to help users better organize and manage the video conference. Moreover, the duration of each session is determined in an automatic manner, which reduces the need for human intervention and improves the efficiency of conference organization and management.

[0017] After the video conference starts, it may be detected whether the video conference proceeds to one of the plurality of sessions. When it is detected that the video conference proceeds to a session, a timer corresponding to the session may be started. The timer may display a remaining duration associated with a planned duration of the session. When the remaining duration of the timer is lower than a predetermined threshold, an indication may be output. The indication may be used to remind a speaker of the session that it is about to time out or it has timed out. The indication may include an audio signal and / or a visual effect. The intensity7of the indication may vary according to the difference between an actual duration of the session and theplanned duration of the session. The larger the difference, that is, the more the actual duration exceeds the planned duration, the greater the intensity of the indication. The speaker may reduce or eliminate the indication through interacting with the indication. The technical effect of the approach described above is to explicitly and intuitively prompt the speaker to control the speaking time, so as to enable the plurality of sessions of the video conference to proceed as planned, while also avoiding abrupt interruptions to the speaker.

[0018] Various embodiments of the present disclosure will hereinafter be described in detail in connection with the appended drawings.

[0019] FIG. 1 illustrates an exemplary process 100 for time analysis in video conference according to an embodiment of the present disclosure. The process 100 may be performed by a video conference system or a video conference application.

[0020] At 102, a total duration of a video conference may be obtained. The total duration may be preset by a host of the video conference.

[0021] At 104, a conference description of the video conference and / or a document associated with the video conference may be obtained. The conference description may be extracted from a conference invitation email of the video conference, w hich may include a set of topics planned to be discussed in the video conference, a set of speakers to speak during the video conference, etc. The document associated with the video conference may be a document attached to the conference invitation email of the video conference, a document shared during the video conference, etc. The document may be various electronic documents processed using document authoring or editing software, including, e.g., a Word processing document (e.g., Word document), a presentation (e.g., a PowerPoint document), a document in Portable Document Format (PDF), etc.

[0022] Subsequently, a plurality of sessions included in the video conference and a planned duration of each session may be determined based on at least one of the total duration, the conference description, and the document.

[0023] Optionally, at 106, a planned duration of an opening session and / or a planned duration of an ending session of the video conference may be determined. The planned duration of the opening session may be set by the host of the video conference or the video conference application according to an empirical value. For example, the planned duration of the opening session may be set to 2 minutes. Alternatively, the planned duration of the opening session may be automatically calculated based on the total duration and a first predetermined ratio. As an example, the first predetermined ratio may be 5%. The planned duration of the ending session may be set by the host of the video conference or the video conference application according to an empirical value. For example, the planned duration of the ending session may be set to 10 minutes. Alternatively, the planned duration of the ending session may be automatically calculated based on the totalduration and a second predetermined ratio. As an example, the second predetermined ratio may be 10%. It should be appreciated that the video conference may not include the opening session or the ending session. Accordingly , the operation at 106 may not be performed.

[0024] A set of content sessions included in the video conference and a planned duration of each content session may be determined in a variety of ways. As previously described, the content session is a session for elaborating or discussing the topic of the conference. A content session may correspond to a topic and / or a speaker.

[0025] In an implementation, the set of content sessions included in the video conference and the planned duration of each content session may be determined by the host of the video conference, as shown in a step 108. For example, an interactive user interface may be presented to the host of the video conference. The host may input a set of content sessions and the planned duration of each content session via the user interface.

[0026] In another implementation, the set of content sessions included in the video conference and the planned duration of each content session may be automatically determined with the conference description of the video conference, as shown in a step 110. The set of content sessions of the video conference and the planned duration of each content session may be determined based on the total duration of the video conference and the conference description through a large language model. An exemplary process for determining the set of content sessions included in the video conference and the planned duration of each content session with the conference description of the video conference will be described below in conjunction with FIG. 2. Preferably, after automatically determining the set of content sessions included in the video conference and the planned duration of each content session with the conference description of the video conference, the determined content sessions and planned durations may be provided to the host of the video conference via email or other means. The host may modify or confirm the content sessions and planned durations.

[0027] In yet another implementation, the set of content sessions included in the video conference and the planned duration of each content session may be automatically determined with the document associated with the video conference, as shown in a step 112. The document includes a plurality of pages, and a duration of each page may be determined based on content of the page. The set of content sessions may be determined through analyzing a set of topics included in the document. For each content session, a set of pages belonging to the content session may be identified, and a planned duration of the content session may be calculated based on a set of durations corresponding to the set of pages. An exemplary process for determining the set of content sessions included in the video conference and the planned duration of each content session with the document associated with the video conference will be described below in conjunctionwith FIG. 3. Preferably, after automatically determining the set of content sessions included in the video conference and the planned duration of each content session with the document associated with the video conference, the determined content sessions and planned durations may be provided to the host of the video conference via email or other means. The host may modify or confirm the content sessions and planned durations.

[0028] The technical effects of the implementations described above are that the duration of each session of the video conference can be reasonably allocated, to help users better organize and manage the video conference. Moreover, the duration of each session is determined in an automatic manner, which reduces the need for human intervention and improves the efficiency of conference organization and management. The set of content sessions included in the video conference and the planned duration of each content session may be determined through one of the plurality7of implementations described above. Preferably, different priorities may be set for the implementations described above. For example, the implementation shown in the step 108 may have the first priority, the implementation shown in the step 110 may have the second priority, and the implementation shown in the step 112 may have the third priority7.

[0029] After the set of content sessions included in the video conference and the planned duration of each content session is determined through one of the plurality of implementations described above, at 114, it may be detected whether the video conference proceeds to one of the plurality of sessions of the video conference. The plurality of sessions may include the opening session, the content sessions, and the ending session. The detection at 114 may be performed in a variety of ways. In an implementation, for each session, it may be detected whether a speaker corresponding to the session begins to speak. In another implementation, for each session, it may be detected whether the document associated with the video conference is switched to a starting page in a set of pages corresponding to the session.

[0030] If it is detected at 114 that the video conference proceeds to the session, the process 100 may proceed to a step 116. At 116, a timer corresponding to the session may be started. The timer may display a remaining duration associated with the planned duration of the session. For example, assuming that the planned duration of the session is 5 minutes, an initial duration of the timer may be 5 minutes. As the video conference progresses, the remaining duration displayed by the timer may be decremented in real time at a predetermined time interval, such as 1 second.

[0031] At 118. it may be determined whether the remaining duration of the timer is below a predetermined threshold. The predetermined threshold may be greater than zero, such as 30 seconds. Alternatively, the predetermined threshold may be equal to zero.

[0032] If it is determined at 118 that the remaining duration of the timer is below7the predetermined threshold, the process 100 may proceed to a step 120. At 120, an indication maybe output. When the predetermined threshold is greater than zero, the indication may be used to remind the speaker of the session that it is about to time out. When the predetermined threshold is equal to zero, the indication may be used to remind the speaker of the session that it has timed out. The indication may be output only to the speaker of the session, or to all participants of the conference.

[0033] The indication may include an audio signal, a visual effect, etc. The audio signal may be ringtone, music, etc. The visual effect may be static text or picture, dynamic animation, etc. The visual effect may be presented in a video stream corresponding to the speaker’ image of the session, or in a video stream corresponding to a shared document, or in both. Only one of the audio signal and the visual effect may be output, or both the audio signal and the visual effect may be output. The technical effects of reminding the speaker that it is about to time out or it has timed out through the audio signal and / or the visual effect are to explicitly and intuitively prompt the speaker to control the speaking time, so as to enable the plurality of sessions of the video conference to proceed as planned, while also avoiding abrupt interruptions to the speaker.

[0034] Preferably, the indication may be output based on at least one of: current time, current season, attributes of a speaker of the session, and attributes of a background image of the speaker, etc. The current time may be in terms of the time of the day. As an example, when the current time is morning, the audio signal may be a crisp bird chirp; while when the current time is afternoon, the audio signal may be a distant bell. The attributes of the speaker may include, e g., the speaker's age, gender, clothing color, etc. As an example, when the speaker wears dark clothes, the output visual effect may also be dark. The attributes of the background image may include, e.g., the background image’s theme, hue. etc. As an example, when the theme of the speaker's background image is a snowy mountain, the visual effect may be falling snowflakes, gradually accumulating ice cubes, etc. As another example, when the theme of the speaker's background image is autumn, the visual effect may be falling withered leaves. The technical effects of outputting the indication based on the current time, the current season, the attributes of the speaker of the session, the attributes of the speaker's background image, etc., are to enable the output indication to match the scene of the video conference, thereby improving the realism and immersion of users participating in the video conference. Examples of outputting visual effects will be illustrated later in conjunction with FIG. 4 and FIG. 5.

[0035] Preferably, the intensity of the indication may vary according to the difference between an actual duration of the session and the planned duration of the session. The larger the difference, that is, the more the actual duration exceeds the planned duration, the greater the intensity of the indication. When the indication is an audio signal, the audio signal will be more rapid or louder. When the indication is a visual effect, the visual effect will occupy a larger proportion in the userinterface. When the visual effect is presented in the video stream corresponding to the speaker's image, the visual effect may gradually obscure the speaker. When the visual effect is presented in the video stream corresponding to the shared document, the visual effect may gradually obscure the document. The technical effects of varying the intensity of the indication according to the difference between the actual duration and the planned duration are to explicitly remind the speaker of the extent of his / her timeout, so as to encourage him / her to finish speaking as soon as possible.

[0036] At 122, it may be determined whether user interaction with the indication output at 120 is detected. A plurality of interaction modes may be provided, and user interactions may be detected through corresponding approaches. In an implementation, a clickable button may be presented in the user interface of the video conference application. Whether user interaction with the indication is detected may be determined through detecting whether the button is clicked. In another implementation, the speaker's action may be detected by a camera or other sensors to determine whether user interaction with the indication is detected. The action may include, e.g., waving a palm, shaking the head, etc. In yet another implementation, it may be determined whether user interaction with the indication is detected through detecting whether the speaker's mouse moves on the screen. The plurality of implementations described above may be performed independently or in combination with each other.

[0037] If user interaction with the indication is detected at 122, the process 100 may proceed to a step 124. At 124, the indication may be reduced or eliminated. For example, when the indication is an audio signal, the volume of the audio signal may be reduced or the audio signal may be turned off. When the indication is a visual effect, the proportion of the visual effect in the user interface may be reduced. When the speaker's action is detected by a camera or other sensors, the area in the visual effect that overlaps with the speaker's joints or other body parts may be detected, and the visual effect of the area may be eliminated, or the visual effect may be retracted below the area. Such interactive effects may be achieved through a physical engine. When the user interaction with the visual effect is detected through detecting the speaker's mouse moving on the screen, the area in the visual effect that overlaps with the mouse passing area may be detected, and the visual effect of the area may be eliminated, or the visual effect may be retracted below7the area. It should be appreciated that, depending on different visual effects or different user interaction modes, the indication may be reduced or eliminated in different ways.

[0038] It should be appreciated that the process for time analysis in video conference described above in conjunction with FIG. 1 is merely exemplary'. Depending on actual application requirements, the steps in the process for time analysis in video conference may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in theprocess 100, the planned duration of the opening session and the planned duration of the ending session are set by the host of the video conference or the video conference application according to the empirical value, or are automatically calculated according to the total duration of the video conference and the predetermined ratios, but the embodiments of the present disclosure are not limited to this. In some embodiments, the planned duration of the opening session and the planned duration of the ending session may also be calculated through a large language model. In addition, the specific order or hierarchy of the steps in the process 100 is merely exemplary7, and the process for time analysis in video conference may be performed in an order different from the described order.

[0039] FIG. 2 illustrates an exemplary7process 200 for determining a set of content sessions included in a video conference and a planned duration of each content session with a conference description of the video conference according to an embodiment of the present disclosure. The process 200 may correspond to the step 110 in FIG. 1. In the process 200, a set of content sessions 232 of a video conference and a planned duration 234 of each content session may be determined, through a large language model 230, based on a total duration 202 of the video conference and a conference description 204 of the video conference.

[0040] The total duration 202 may be preset by a host of the video conference. The conference description 204 may be extracted from a conference invitation email of the video conference, which may include a set of topics planned to be discussed in the video conference, a set of speakers to speak during the video conference, etc.

[0041] A prompt 222 to be provided to the large language model 230 may be constructed through a prompt constructor 220. The prompt constructor 220 may construct the prompt 222 based on the total duration 202 and the conference description 204. Preferably, a prompt template 210 for the prompt constructor 220 may be designed in advance. The prompt template 210 may include a plurality7of variable parts for loading the total duration 202 and the conference description 204, respectively. In addition, the prompt template 210 may include a response instruction for instructing the large language model 230 on how to respond. Optionally, in the case where the video conference includes an opening session and / or an ending session, when constructing the prompt 222, a planned duration 206 of the opening session and / or a planned session 208 of the ending session may also be considered accordingly.

[0042] The large language model 230 is, e.g., a Generative Pre-trained Transformer (GPT) model. The large language model 230 may identify, from the conference description 204 included in the prompt 222, a set of topics planned to be discussed in the video conference, a set of speakers to speak during the video conference, etc. with its logical inference ability, big data support ability, etc., thereby determining a set of content sessions 232 included in the video conference. Inaddition, the large language model 230 may determine the duration allocated to each topic or each speaker based on the total duration 202, thereby determining the planned duration 234 of each content session. In the case where the prompt 222 includes the planned duration 206 of the opening session and / or the planned duration 208 of the ending session, the large language model 230 may also determine the planned duration 234 of each content session based on the planned duration 206 of the opening session and / or the planned duration 208 of the ending session. Preferably, the large language model 230 may determine several idle time periods reserved for switches between sessions or additional discussions.

[0043] It should be appreciated that the process for determining the set of content sessions included in the video conference and the planned duration of each content session with the conference description of the video conference described above in conjunction with FIG. 2 is merely exemplary. Depending on actual application requirements, the steps in the process for determining the set of content sessions included in the video conference and the planned duration of each content session with the conference description of the video conference may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in the process 200, the large language model 230 is used to determine the planned duration of each content session, but the embodiments of the present disclosure are not limited to this. In some embodiments, a calculation unit capable of performing mathematical operations may be used to determine the planned duration. For example, the number of topics or the number of speakers included in the conference description 204 may be identified, so as to determine the number of content sessions. The duration reserved for the set of content sessions may be calculated based on the total duration 202, the planned duration 206 of the opening session, and / or the planned duration 208 of the ending session. The duration may be evenly distributed to the set of content sessions, so as to determine the planned duration of each content session.

[0044] FIG. 3 illustrates an exemplary' process 300 for determining a set of content sessions included in a video conference and a planned duration of each content session with a document associated with the video conference according to an embodiment of the present disclosure. The process 300 may correspond to the step 112 in FIG. 1.

[0045] The document associated with the video conference may be a document attached to a conference invitation email of the video conference, a document shared during the video conference, etc. The document may be various electronic documents processed using document authoring or editing software, including, e.g., Word document, PowerPoint document, PDF document, etc.

[0046] At 302, for each page in a plurality' of pages of the document, a duration of the page may be determined based on content of the page. This step may be performed by a large languagemodel. Under the guidance of a prompt provided to the large language model, the large language model may determine a duration of each page based on content of the page, such as the number of words in the text, the duration of video or audio playback, etc. The duration of the page pLmay be denoted as t(p . The technical effects of determining the duration of the page based on the content of the page are to cause the determined duration of the page to be related to the complexity of the page, so that the determined duration is accurate. Preferably, the prompt provided to the large language model may also include the total duration of the video conference. In the case where the video conference includes an opening session and / or an ending session, the prompt may also include a planned duration of the opening session and / or a planned duration of the ending session. The technical effect of including the total duration, the planned duration of the opening session and / or the planned duration of the ending session in the prompt is to cause the duration of each page determined through the large language model to be within a reasonable time range.

[0047] At 304. a set of content sessions of the video conference may be determined through analyzing a set of topics included in the document. This step may be performed by a large language model. Under the guidance of a prompt provided to the large language model, the large language model may identify a set of topics included in the document through analyzing a director}' page or a conference agenda page of the document. Preferably, when a plurality of topics correspond to a same speaker, a plurality of content sessions corresponding to the plurality of topics may be merged into one content session. The technical effects of this approach are to reduce the number of indications for the same speaker, so as to reduce the frequency of interfering with the speaker.

[0048] At 306, for each content session in the set of content sessions, a set of pages in the plurality of pages of the document belonging to a topic corresponding to the content session may be identified. For example, each page may be classified into one of the set of topics identified at 304 based on title, content, etc. of the page, thereby identifying a set of pages corresponding to each content session. This operation may be performed by a large language model. An initial planned duration of the content session may be calculated based on a set of durations corresponding to the set of pages. For example, the initial planned duration session session^ may be calculated through the following equation:

[0049] At 308, a free duration fyreemay be calculated based on the total duration Ttotaiof the video conference and a set of initial planned durations [77^-“^.} corresponding to the set of content sessions, as shown in the following equation:Where N represents the number of content sessions of the video conference.

[0050] Preferably, in the case where the video conference includes an opening session and / or an ending session, the free duration Tfreemay also be calculated based on a planned duration Topening of the opening session and / or a planned duration Tendingof the ending session, as shown in the following equation:

[0051] At 310, an additional duration of each content session may be calculated based on the free duration. The additional duration of the content session session, may be denoted as..

[0052] In an implementation, the additional durations of all content sessions may be equal. The free duration T^reemay be evenly distributed to each content session according to the number N of content sessions. For example, the additional duration T°ddnn. of the content session session.] may be calculated through the following equation:

[0053] In another implementation, the additional duration of the content session may be related to the ratio of the initial planned duration of the content session to the total initial planned duration. The longer the initial planned duration, the longer its additional duration. For example, the additional duration of the content session session J, mav be calculated through thefollowing equation:

[0054] At 312, for each content session in the set of content sessions, a planned duration of the content session may be calculated based on the initial planned duration of the content session and the additional duration of the content session. The planned duration. of the content session session] may be calculated, e.g., through the following equation:

[0055] It should be appreciated that the process for determining the set of content sessions included in the video conference and the planned duration of each content session with the document associated with the video conference described above in conjunction with FIG. 3 is merely exemplary. Depending on actual application requirements, the steps in the process fordetermining the set of content sessions and the planned duration of each content session with the document may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in the process 300, after the initial planned duration of the content session is calculated through the step 306, the additional duration of the content session is also calculated through the step 308 and the step 310, and the planned duration of the content session is calculated based on the initial planned duration and the additional duration at the step 312, but the embodiments of the present disclosure are not limited to this. In some embodiments, the step 308 to the step 312 may be omitted. In this case, the initial planned duration obtained through the step 306 may be directly used as the final planned duration. In addition, the specific order or hierarchy of the steps in the process 300 is merely exemplary, and the process for determining the set of content sessions and the planned duration of each content session with the document may be performed in an order different from the described order.

[0056] FIG. 4 illustrates an example of outputting a visual effect according to an embodiment of the present disclosure. Diagram 400a is, e.g., a video frame in a video stream corresponding to an image of a speaker 402. The speaker 402 is, e.g., a speaker of a session of a video conference. The session may be an opening session, a content session, or an ending session. A background image 404 of the speaker 402 is a virtual background image. The background image 404 includes snow-capped mountains, white clouds, etc. A timer 406 is shown at the top of the diagram 400a. The timer 406 displays a remaining duration associated with a planned duration of the session, e.g., 3 minutes and 45 seconds.

[0057] As the video conference progresses, the remaining duration displayed by the timer 406 may be decremented in real time at a predetermined time interval, such as 1 second. According to the embodiments of the present disclosure, when the remaining duration is lower than a predetermined threshold, an indication may be output to remind the speaker 402 that it is about to time out or it has timed out. The indication may be an audio signal and / or a visual effect. According to the embodiments of the present disclosure, the indication may be output based on at least one of: current time, current season, attributes of a speaker of the session, and attributes of a background image of the speaker, etc.

[0058] An example of outputting a visual effect is show n in a diagram 400b. In the diagram 400b, the remaining duration displayed by the timer 406 is minus 35 seconds, which means that the planned duration has been exceeded by 35 seconds. Compared with the timer 406 in the diagram 400a, the color of the timer 406 in the diagram 400b has also changed, becoming more eye-catching, which can more clearly remind the speaker that it has timed out. Since the background image 404 includes a snowy mountain, when the remaining duration displayed by the timer 406 is lower than a predetermined threshold, an output visual effect 408 may be fallingsnowflakes. The snowflakes may fall on and near the speaker 402.

[0059] According to the embodiments of the present disclosure, the intensity of the visual effect may vary' according to the difference between an actual duration and the planned duration of the session. The larger the difference, that is, the more the actual duration exceeds the planned duration, the greater the intensity of the visual effect. For example, assuming that as time passes, the speaker 402 still has not finished speaking, the snowflakes falling on and near the speaker 402 will increase, and even completely cover the speaker 402.

[0060] The speaker 402 may reduce or eliminate the visual effect 408 through interacting with the visual effect 408. For example, the speaker 402 may interact with the visual effect 408 through moving a mouse on the screen, waving a palm, shaking the head, or performing other actions. If user interaction with the visual effect 408 is detected, the proportion of the visual effect 408 in the user interface may be reduced, such as reducing the number of snowflakes, reducing the speed at which the snowflakes fall, etc.

[0061] FIG. 5 illustrates another example of outputting a visual effect according to an embodiment of the present disclosure. Diagram 500a is, e.g., a video frame in a video stream corresponding to an image of a speaker 502. The speaker 502 is, e.g., a speaker of a session of a video conference. The session may be an opening session, a content session, or an ending session. A background image 504 of the speaker 502 is his / her actual background image. A timer 506 is shown at the top of the diagram 500a. The timer 506 displays a remaining duration associated with a planned duration of the session, e.g., 1 minutes and 50 seconds.

[0062] As the video conference progresses, the remaining duration displayed by the timer 506 may be decremented in real time at a predetermined time interval, such as 1 second. According to the embodiments of the present disclosure, when the remaining duration is lower than a predetermined threshold, an indication may be output to remind the speaker 502 that it is about to time out or it has timed out. The indication may be an audio signal and / or a visual effect. According to the embodiments of the present disclosure, the indication may be output based on at least one of: current time, current season, attributes of a speaker of the session, and attributes of a background image of the speaker, etc.

[0063] An example of outputting a visual effect is shown in diagram 500b. In the diagram 500b, the remaining duration displayed by the timer 506 is minus 20 seconds, which means that the planned duration has been exceeded by 20 seconds. Compared with the timer 506 in the diagram 500a, the color of the timer 506 in the diagram 500b has also changed, becoming more eye-catching, which can more clearly remind the speaker that it has timed out. A visual effect 508 shown in the diagram 500b is a gradually closing curtain. Since the speaker 502 is wearing dark clothing, the color of the curtain is also dark, or the same as the color of the speaker's 502 clothing.

[0064] According to the embodiments of the present disclosure, the intensity of the visual effect may vary according to the difference between an actual duration and the planned duration of the session. The larger the difference, that is, the more the actual duration exceeds the planned duration, the greater the intensity of the visual effect. For example, assuming that as time passes, the speaker 502 still has not finished speaking, the curtain may be gradually closed, or even completely cover the speaker 502.

[0065] The speaker 502 may reduce or eliminate the visual effect 508 through interacting with the visual effect 508. For example, the speaker 502 may interact with the visual effect 508 through moving a mouse on the screen, waving a palm, shaking the head, or performing other actions. If user interaction with the visual effect 508 is detected, the proportion of the visual effect 508 in the user interface may be reduced, e.g., making the closing curtain open to both sides, reducing the closing speed of the curtain, etc.

[0066] It should be appreciated that the visual effects shown in FIG. 4 and FIG. 5 are merely two examples of visual effects. Depending on actual application requirements, other visual effects may also be employed to remind the speaker that it is about to time out or it has timed out.

[0067] FIG. 6 is a flowchart of an exemplary' method 600 for time analysis in video conference according to an embodiment of the present disclosure.

[0068] At 610. a total duration of a video conference may be obtained.

[0069] At 620, a conference description of the video conference and / or a document associated with the video conference may be obtained.

[0070] At 630, a plurality of sessions included in the video conference and a planned duration of each session may be determined based on at least one of the total duration, the conference description, and the document.

[0071] At 640, for each session in the plurality of sessions, it may be detected that the video conference proceeds to the session.

[0072] At 650, in response to detecting that the video conference proceeds to the session, a timer corresponding to the session may be started, the timer displaying a remaining duration associated with a planned duration of the session.

[0073] At 660, in response to the remaining duration being below a predetermined threshold, an indication may be output, the indication including an audio signal and / or a visual effect.

[0074] In an implementation, the plurality of sessions may include an opening session. A planned duration of the opening session may be calculated based on the total duration and a first predetermined ratio. The plurality of sessions may include an ending session. A planned duration of the end session may be calculated based on the total duration and a second predetermined ratio.

[0075] In an implementation, the plurality of sessions may include a set of content sessions.Each content session may correspond to a topic and / or a speaker. The determining a plurality of sessions included in the video conference and a planned duration of each session may comprise: determining the set of content sessions and a planned duration of each content session based on the total duration and the conference description through a large language model.

[0076] In an implementation, the plurality of sessions may include a set of content sessions. Each content session corresponds to a topic and / or a speaker. The determining a plurality of sessions included in the video conference and a planned duration of each session may comprise: for each page in a plurality of pages of the document, determining a duration of the page based on content of the page; determining the set of content sessions through analyzing a set of topics included in the document; for each content session in the set of content sessions, identifying a set of pages in the plurality of pages belonging to a topic corresponding to the content session, and calculating an initial planned duration of the content session based on a set of durations corresponding to the set of pages; calculating a free duration based on the total duration and a set of initial planned durations corresponding to the set of content sessions; calculating an additional duration of each content session based on the free duration; and for each content session in the set of content sessions, calculating a planned duration of the content session based on the initial planned duration of the content session and the additional duration of the content session.

[0077] In an implementation, the detecting that the video conference proceeds to the session may comprise: detecting that a speaker corresponding to the session begins to speak; and / or detecting that the document is switched to a starting page in a set of pages corresponding to the session.

[0078] In an implementation, the outputting an indication may comprise: outputting the indication based on at least one of: current time, current season, attributes of a speaker of the session, and attributes of a background image of the speaker.

[0079] In an implementation, the intensify' of the indication may vary according to the difference between an actual duration of the session and the planned duration of the session.

[0080] In an implementation, the method 600 may further comprise: detecting user interaction with the indication; and in response to detecting the user interaction with the indication, reducing or eliminating the indication.

[0081] It should be appreciated that the method 600 may further comprise any other steps / processes for time analysis in video conference according to the embodiments of the present disclosure as mentioned above.

[0082] FIG. 7 illustrates an exemplary apparatus 700 for time analysis in video conference according to an embodiment of the present disclosure.

[0083] The apparatus 700 may comprise: a total duration obtaining module 710, for obtaininga total duration of a video conference; a conference description and / or document obtaining module 720, for obtaining a conference description of the video conference and / or a document associated with the video conference; a session and planned duration determining module 730, for determining a plurality of sessions included in the video conference and a planned duration of each session based on at least one of the total duration, the conference description, and the document; a session detecting module 740 for, for each session in the plurality of sessions, detecting that the video conference proceeds to the session; a timer starting module 750, for in response to detecting that the video conference proceeds to the session, starting a timer corresponding to the session, the timer displaying a remaining duration associated with a planned duration of the session; and an indication outputting module 760, for in response to the remaining duration being below a predetermined threshold, outputting an indication, the indication including an audio signal and / or a visual effect. Furthermore, the apparatus 700 may further comprise any other modules configured for time analysis in video conference according to the embodiments of the present disclosure as mentioned above.

[0084] FIG. 8 illustrates another exemplary apparatus 800 for time analysis in video conference according to an embodiment of the present disclosure.

[0085] The apparatus 800 may comprise a processor 810; and a memory' 820 storing computer-executable instructions. The computer executable instructions, when executed, may cause the processor 810 to: obtain a total duration of a video conference; obtain a conference description of the video conference and / or a document associated with the video conference; determine a plurality of sessions included in the video conference and a planned duration of each session based on at least one of the total duration, the conference description, and the document; and for each session in the plurality of sessions: detect that the video conference proceeds to the session; in response to detecting that the video conference proceeds to the session, start a timer corresponding to the session, the timer displaying a remaining duration associated with a planned duration of the session; and in response to the remaining duration being below a predetermined threshold, output an indication, the indication including an audio signal and / or a visual effect.

[0086] In an implementation, the plurality of sessions may include an opening session. A planned duration of the opening session may be calculated based on the total duration and a first predetermined ratio. The plurality of sessions may include an ending session. A planned duration of the end session may be calculated based on the total duration and a second predetermined ratio.

[0087] In an implementation, the plurality of sessions may include a set of content sessions. Each content session may correspond to a topic and / or a speaker. The determining a plurality of sessions included in the video conference and a planned duration of each session may comprise: determining the set of content sessions and a planned duration of each content session based onthe total duration and the conference description through a large language model.

[0088] In an implementation, the plurality of sessions may include a set of content sessions. Each content session may correspond to a topic and / or a speaker. The determining a plurality of sessions included in the video conference and a planned duration of each session may comprise: for each page in a plurality of pages of the document, determining a duration of the page based on content of the page; determining the set of content sessions through analyzing a set of topics included in the document; for each content session in the set of content sessions, identifying a set of pages in the plurality of pages belonging to a topic corresponding to the content session, and calculating an initial planned duration of the content session based on a set of durations corresponding to the set of pages; calculating a free duration based on the total duration and a set of initial planned durations corresponding to the set of content sessions; calculating an additional duration of each content session based on the free duration; and for each content session in the set of content sessions, calculating a planned duration of the content session based on the initial planned duration of the content session and the additional duration of the content session.

[0089] In an implementation, the detecting that the video conference proceeds to the session may comprise: detecting that a speaker corresponding to the session begins to speak; and / or detecting that the document is switched to a starting page in a set of pages corresponding to the session.

[0090] In an implementation, the outputting an indication may comprise: outputting the indication based on at least one of: current time, current season, attributes of a speaker of the session, and attributes of a background image of the speaker.

[0091] In an implementation, the intensity of the indication may vary according to the difference between an actual duration of the session and the planned duration of the session.

[0092] In an implementation, the computer executable instructions, when executed, may further cause the processor 810 to: detect user interaction wi th the indication; and in response to detecting the user interaction with the indication, reduce or eliminate the indication.

[0093] It should be appreciated that the processor 810 may further perform any other steps / processes of the method for time analysis in video conference according to the embodiments of the present disclosure as mentioned above.

[0094] The embodiments of the present disclosures propose a computer program product for time analysis in video conference, comprising a computer program that is executed by a processor for: obtaining a total duration of a video conference; obtaining a conference description of the video conference and / or a document associated with the video conference; determining a plurality of sessions included in the video conference and a planned duration of each session based on at least one of the total duration, the conference description, and the document; and for each sessionin the plurality of sessions: detecting that the video conference proceeds to the session; in response to detecting that the video conference proceeds to the session, starting a timer corresponding to the session, the timer displaying a remaining duration associated with a planned duration of the session; and in response to the remaining duration being below a predetermined threshold, outputting an indication, the indication including an audio signal and / or a visual effect. Furthermore, the computer program may be further executed for implementing any other steps / processes of the method for time analysis in video conference according to the embodiments of the present disclosure as mentioned above.

[0095] The embodiments of the present disclosure may be embodied in a computer-readable medium for time analysis in video conference. The computer-readable medium may comprise instructions that, when executed, cause a processor to: obtain a total duration of a video conference; obtain a conference description of the video conference and / or a document associated with the video conference; determine a plurality’ of sessions included in the video conference and a planned duration of each session based on at least one of the total duration, the conference description, and the document; and for each session in the plurality of sessions: detect that the video conference proceeds to the session; in response to detecting that the video conference proceeds to the session, start a timer corresponding to the session, the timer displaying a remaining duration associated with a planned duration of the session; and in response to the remaining duration being below a predetermined threshold, output an indication, the indication including an audio signal and / or a visual effect. Furthermore, the instructions, when executed, may further cause the processor to perform any other steps / processes of the method for time analysis in video conference according to the embodiments of the present disclosure as mentioned above.

[0096] It should be appreciated that all the operations in the methods described above are merely exemplary, and the present disclosure is not limited to any operations in the methods or sequence orders of these operations, and should cover all other equivalents under the same or similar concepts. In addition, the articles “a” and "an" as used in this specification and the appended claims should generally be construed to mean “one” or “one or more” unless specified otherwise or clear from the context to be directed to a singular form.

[0097] It should also be appreciated that all the modules in the apparatuses described above may be implemented in various approaches. These modules may be implemented as hardware, software, or a combination thereof. Moreover, any of these modules may be further functionally divided into sub-modules or combined together.

[0098] Processors have been described in connection with various apparatuses and methods. These processors may be implemented using electronic hardware, computer software, or any combination thereof. Whether such processors are implemented as hardware or software willdepend upon the particular application and overall design constraints imposed on the system. By way of example, a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with a microprocessor, microcontroller, digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD). a state machine, gated logic, discrete hardware circuits, and other suitable processing components configured for performing the various functions described throughout the present disclosure. The functionality of a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with software being executed by a microprocessor, microcontroller, DSP, or other suitable platform.

[0099] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, threads of execution, procedures, functions, etc. The software may reside on a computer-readable medium. A computer-readable medium may include, by way of example, memory such as a magnetic storage device (e.g., hard disk, floppy disk, magnetic strip), an optical disk, a smart card, a flash memory device, random access memory7(RAM), read only memory7(ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), a register, or a removable disk. Although memory is shown separate from the processors in the various aspects presented throughout the present disclosure, the memory may be internal to the processors, e g., cache or register.

[0100] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described throughout the present disclosure that are known or later come to be known to those of ordinary7skilled in the art are expressly incorporated herein and intended to be encompassed by the claims.

Claims

CLAIMS1 . A method for time analysis in video conference, comprising: obtaining a total duration of a video conference; obtaining a conference description of the video conference and / or a document associated with the video conference; determining a plurality of sessions included in the video conference and a planned duration of each session based on at least one of the total duration, the conference description, and the document; and for each session in the plurality of sessions: detecting that the video conference proceeds to the session; in response to detecting that the video conference proceeds to the session, starting a timer corresponding to the session, the timer displaying a remaining duration associated with a planned duration of the session; and in response to the remaining duration being below a predetermined threshold, outputting an indication, the indication including an audio signal and / or a visual effect.

2. The method of claim 1, wherein: the plurality of sessions include an opening session, and a planned duration of the opening session is calculated based on the total duration and a first predetermined ratio, and / or the plurality of sessions include an ending session, and a planned duration of the end session is calculated based on the total duration and a second predetermined ratio.

3. The method of claim 1, wherein the plurality of sessions include a set of content sessions, each content session corresponding to a topic and / or a speaker, and the determining a plurality of sessions included in the video conference and a planned duration of each session comprises: determining the set of content sessions and a planned duration of each content session based on the total duration and the conference description through a large language model.

4. The method of claim 1, wherein the plurality of sessions include a set of content sessions, each content session corresponding to a topic and / or a speaker, and the determining a plurality of sessions included in the video conference and a planned duration of each session comprises: for each page in a plurality of pages of the document, determining a duration of the page based on content of the page; determining the set of content sessions through analyzing a set of topics included in the document; for each content session in the set of content sessions, identifying a set of pages in the plurality of pages belonging to a topic corresponding to the content session, and calculating an initial planned duration of the content session based on a set of durations corresponding to the setof pages; calculating a free duration based on the total duration and a set of initial planned durations corresponding to the set of content sessions; calculating an additional duration of each content session based on the free duration; and for each content session in the set of content sessions, calculating a planned duration of the content session based on the initial planned duration of the content session and the additional duration of the content session.

5. The method of claim 1, wherein the detecting that the video conference proceeds to the session comprises: detecting that a speaker corresponding to the session begins to speak; and / or detecting that the document is switched to a starting page in a set of pages corresponding to the session.

6. The method of claim 1, wherein the outputting an indication comprises: outputting the indication based on at least one of: cunent time, current season, attributes of a speaker of the session, and attributes of a background image of the speaker.

7. The method of claim 1, wherein the intensity of the indication varies according to the difference betw een an actual duration of the session and the planned duration of the session.

8. The method of claim 1, further comprising: detecting user interaction with the indication; and in response to detecting the user interaction with the indication, reducing or eliminating the indication.

9. An apparatus for time analysis in video conference, comprising: a processor; and a memory storing computer-executable instructions that, when executed, cause the processor to: obtain a total duration of a video conference; obtain a conference description of the video conference and / or a document associated with the video conference; determine a plurality of sessions included in the video conference and a planned duration of each session based on at least one of the total duration, the conference description, and the document; and for each session in the plurality of sessions: detect that the video conference proceeds to the session; in response to detecting that the video conference proceeds to the session, start a timer corresponding to the session, the timer displaying a remaining duration associatedwith a planned duration of the session; and in response to the remaining duration being below a predetermined threshold, output an indication, the indication including an audio signal and / or a visual effect.

10. The apparatus of claim 9, wherein: the plurality of sessions include an opening session, and a planned duration of the opening session is calculated based on the total duration and a first predetermined ratio, and / or the plurality of sessions include an ending session, and a planned duration of the end session is calculated based on the total duration and a second predetermined ratio.

11. The apparatus of claim 9, wherein the plurality of sessions include a set of content sessions, each content session corresponding to a topic and / or a speaker, and the determining a plurality of sessions included in the video conference and a planned duration of each session comprises: determining the set of content sessions and a planned duration of each content session based on the total duration and the conference description through a large language model.

12. The apparatus of claim 9, wherein the plurality of sessions include a set of content sessions, each content session corresponding to a topic and / or a speaker, and the determining a plurality of sessions included in the video conference and a planned duration of each session comprises: for each page in a plurality of pages of the document, determining a duration of the page based on content of the page; determining the set of content sessions through analyzing a set of topics included in the document; for each content session in the set of content sessions, identifying a set of pages in the plurality of pages belonging to a topic corresponding to the content session, and calculating an initial planned duration of the content session based on a set of durations corresponding to the set of pages; calculating a free duration based on the total duration and a set of initial planned durations corresponding to the set of content sessions; calculating an additional duration of each content session based on the free duration; and for each content session in the set of content sessions, calculating a planned duration of the content session based on the initial planned duration of the content session and the additional duration of the content session.

13. The apparatus of claim 9, wherein the intensity of the indication varies according to the difference between an actual duration of the session and the planned duration of the session.

14. The apparatus of claim 9. wherein the computer-executable instructions, when executed,further cause the processor to: detect user interaction with the indication; and in response to detecting the user interaction with the indication, reduce or eliminate the indication.

15. A computer program product for time analysis in video conference, comprising a computer program that is executed by a processor for: obtaining a total duration of a video conference; obtaining a conference description of the video conference and / or a document associated with the video conference; determining a plurality of sessions included in the video conference and a planned duration of each session based on at least one of the total duration, the conference description, and the document; and for each session in the plurality of sessions: detecting that the video conference proceeds to the session; in response to detecting that the video conference proceeds to the session, starting a timer corresponding to the session, the timer displaying a remaining duration associated with a planned duration of the session; and in response to the remaining duration being below a predetermined threshold, outputting an indication, the indication including an audio signal and / or a visual effect.

Citation Information

Patent Citations

  • Advising meeting participants of their contributions based on a graphical representation

    EP3783553A1