Device, estimation method, and program

A device and method for estimating user attention in multi-user video systems by analyzing screen size and audio/video states address the limitations of existing technologies, providing accurate attention level estimation and quality feedback.

WO2025210712A1PCT designated stage Publication Date: 2025-10-09NT T INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/013512
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-01
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing technologies lack a method to accurately estimate user attention levels in systems displaying multiple video streams, such as web conferences, due to reliance on single-image/video analysis, high computational costs, and neglect of audio and temporal aspects.

Method used

A device and method that estimate user attention levels by analyzing screen display size and audio/video state parameters, including microphone and camera usage, across multiple user terminals, integrating spatial and temporal attention metrics.

Benefits of technology

Enables accurate estimation of user attention levels and quality of experience in multi-user video systems, improving communication strategies and feedback for presentations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024013512_09102025_PF_FP_ABST
    Figure JP2024013512_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A device, used in a system in which a screen in which images of one or more users are arranged is displayed on terminals of each of a plurality of users, comprises an estimation unit for estimating the level of attention given to each of the users by a specific user, the attention level being estimated using either or both the screen display size of the images of each of the users and a state parameter of each of the users on the screen of the specific user.
Need to check novelty before this filing date? Find Prior Art

Description

Apparatus, estimation method, and program The present invention relates to an estimation technique in a system in which multiple videos are displayed on a screen, such as a web conference service. In a Web conference service, some participants often receive attention and others do not. If we can estimate the level of attention, we can collect and analyze data after the meeting on which speakers attracted the most attention and use this data to create feedback to improve future presentations and discussion styles. In addition, it can be used to understand which participants are most influential during a meeting and which participants are not participating, which can be used to understand participant groups, and to analyze interaction patterns between participants and help develop strategies to improve communication imbalances within a team.Furthermore, since the video quality and audio quality of a highly visible participant are considered to have a greater impact than the quality of other participants, it can also be used to estimate the quality of a web conference. Ivan Himawan et al.

[2013] ,"Automatic Region-of-Interest Detection and Prioritization for Visually Optimized Coding of Low Bit Rate Videos", IEEE Workshop on Applications of Computer Vision (WACV) However, no technology has been established for estimating attention levels from easily obtainable parameters. Non-Patent Document 1 discloses a technology for estimating a region of interest (ROI) in an image or video. However, the conventional technology disclosed in Non-Patent Document 1 is based on the premise of a single image or video, and is therefore not suitable for web conferences in which videos of multiple participants are displayed on the screen. In addition, most of the conventional techniques analyze image / video signals, which requires a large amount of calculation, resulting in high costs.Furthermore, since they do not take into account audio or temporal aspects, it is difficult to achieve highly accurate estimation. The above-mentioned problem is not limited to video images in a web conference, but can occur in any system in which the videos of multiple users are displayed on the screens of the terminals. The present invention has been made in consideration of the above points, and aims to provide a technology for estimating the level of attention of a user in a system in which videos of multiple users are displayed on the screens of each terminal. According to the disclosed technology, there is provided a device used in a system in which a screen on which one or more user's images are arranged is displayed on each of a plurality of user's terminals, the device comprising: an estimation unit that estimates the degree of attention of the specific user to each user by using either or both of the screen display size of the video of each user on the specific user's screen and the state parameters of each user; An apparatus is provided comprising: The disclosed technology provides a technology for estimating the attention level of a user in a system in which videos of multiple users are displayed on the screens of each terminal. FIG. 1 is a configuration diagram of a communication system. FIG. 2 is a diagram showing an example of a screen. FIG. 3 is a configuration diagram of an attention level estimation device 100. FIG. 4 is a flowchart showing the operation of the attention level estimation device 100. FIG. 5 is a diagram showing a specific example of Example 1. FIG. 6 is a diagram showing a specific example of Example 1. FIG. 7 is a diagram showing a specific example of Example 2. FIG. 8 is a diagram showing a specific example of Example 3. FIG. 9 is a diagram showing a specific example of Example 3. FIG. 10 is a diagram showing a specific example of Example 4. FIG. 11 is a diagram showing an example of the hardware configuration of an apparatus. Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment. Note that, below, we will explain a technology for estimating the attention level of each participant in a web conferencing service, but the technology according to this embodiment is not limited to web conferencing services and can be applied to any system in which videos of multiple users are displayed on the screen of each terminal. (Outline of the embodiment) In this embodiment, an attention level estimation device 100 (described later) estimates the level of attention a participant pays to other participants by using the screen display size or the behavior of each user (such as a change in mute status). This technology makes it possible to estimate the user quality of experience taking into account, for example, the level of attention each participant pays to other participants. (System configuration example, operation example) Fig. 1 shows an example of the configuration of a communication system according to the present embodiment. As shown in Fig. 1, in this embodiment, a plurality of user terminals 10 are connected to a network 20, and communication is performed between the plurality of user terminals 10 using a web conference service. A web conference application screen (hereinafter referred to as an app screen) is displayed on each user terminal 10. A user of each user terminal 10 (a web conference participant) holds a web conference with one or more other participants while viewing the app screen. 2, for example, images of the multiple participants (e.g., images of their faces) are displayed on each user terminal 10. The app screens displayed on the multiple user terminals 10 may be different. 1 , an attention level estimation device 100 is provided on a network 20. The attention level estimation device 100 can acquire information necessary for attention level estimation from the network 20. Note that the attention level estimation device 100 may be a certain user terminal 10, a server that provides a web conference service, or a device that is neither a user terminal 10 nor the server. The attention level estimation device 100 can estimate, for each user terminal 10 (for each participant), the participant's level of attention to each participant (which may be the participant himself or herself or another person). 3 shows an example of the configuration of the attention level estimation device 100 according to this embodiment. As shown in FIG. 3, the attention level estimation device 100 includes a spatial attention level estimation unit 110, a short-term attention level estimation unit 120, a long-term attention level estimation unit 130, and an attention level integration unit 140. The spatial attention level estimation unit 110, the short-term attention level estimation unit 120, and the long-term attention level estimation unit 130 may be collectively referred to as the "estimation unit." Furthermore, "(at least one of the spatial attention level estimation unit 110, the short-term attention level estimation unit 120, and the long-term attention level estimation unit 130) and the attention level integration unit 140" may be referred to as the "estimation unit." Furthermore, the attention level estimation device 100 may be a device that includes at least one of the spatial attention level estimation unit 110, the short-term attention level estimation unit 120, and the long-term attention level estimation unit 130, but does not include the attention level integration unit 140. Furthermore, the attention level estimation device 100 may be referred to as the "device." The spatial attention estimation unit 110 inputs the screen display size (e.g., area) of each participant for the app screen of a certain participant (let's call it participant A), estimates the attention level based on the proportion that the screen display size occupies in relation to the screen display size of all participants, and outputs the estimated attention level as participant A's spatial attention level for each participant. The short-term attention estimation unit 120 inputs the state parameters of each participant (let's call them participant A) (e.g., the microphone's ON / OFF status, etc.), estimates the attention level based on the proportion of time the microphone was ON within a specific period of time (e.g., one minute), and outputs this attention level as participant A's short-term attention level for each participant. The long-term attention level estimation unit 130 estimates the level of attention of a certain participant (say participant A) based on the state parameters of each participant (e.g., microphone ON / OFF state) and the proportion of time that the microphone was ON within a long period of time (e.g., one hour), and outputs the estimated level of attention as the long-term attention level of participant A for each participant. The short-term attention level and the long-term attention level may be collectively referred to as the time attention level. The attention level integration unit 140 receives at least two of the spatial attention level, short-term attention level, and long-term attention level as input, estimates an attention level by combining the input attention levels, and outputs the estimated attention level value. As described above, the attention level estimation device 100 may not include the attention level integration unit 140. In this case, the output of any one of the spatial attention level estimation unit 110, the short-term attention level estimation unit 120, and the long-term attention level estimation unit 130 may be used as the output of the attention level estimation device 100. Furthermore, using "1 minute" as a "short time" in the short-term attention level is one example, and a time other than "1 minute" may be used. Furthermore, using "1 hour" as a "long time" in the long-term attention level is one example, and a time other than "1 hour" may be used. The "short time" in the short-term attention level may be shorter than the "long time" in the long-term attention level. Furthermore, the state parameters are not limited to those related to audio, but may also be those related to video (for example, the ON / OFF state of a camera). When using the ON / OFF state of a camera, the attention level (short-term attention level, long-term attention level) may be calculated using the same calculation method as when using the ON / OFF state of a microphone, for example. An example of the processing procedure of the attention level estimation device 100 will be described with reference to the flowchart in Fig. 4. Fig. 4 shows a procedure for estimating the level of attention of a certain participant (assumed to be participant A) to a certain participant. The flow in Fig. 4 is executed for each participant, thereby obtaining the level of attention of participant A to each participant. In the example below, three levels of attention are used: spatial attention level, short-term attention level, and long-term attention level. In S101, the spatial attention level estimation unit 110 estimates the spatial attention level, the short-term attention level estimation unit 120 estimates the short-term attention level, and the long-term attention level estimation unit 130 estimates the long-term attention level. In S102, the attention level integration unit 140 combines the spatial attention level, the short-term attention level, and the long-term attention level to estimate and output a final attention level. Detailed processing examples of each part will be described below as Examples 1 to 4. In the following, the subject of interest is the participant P. Example 1 First, a description will be given of Example 1. Example 1 is an example of the spatial attention estimation unit 110. Depending on the web conference application, the user terminal 10 of the participant P can enlarge and display the facial image of a specific participant. The enlarged participant is considered to be of high interest to the participant P. Therefore, the spatial attention level estimation unit 110 estimates the attention level of the specific participant according to the area of ​​the application screen occupied by the facial image of the specific participant. Examples of calculations performed by the spatial attention level estimation unit 110 include the following calculation examples 1-1 to 1-4. In each of the following formulas, A i is the attention level of participant i, and S i is the display area of ​​participant i on the application screen of participant P, and S = Σ i S i is the total participant display area on the app screen of participant P, and N is the number of participants. Calculation example 1-1) Calculation example 1-2) Calculation example 1-3) In the above formula, r is a constant greater than or equal to 0. Calculation example 1-4) In the above equation, b is a constant greater than one. Example 1: Specific Example If there are three participants and the values ​​of each variable on the app screen of participant P (one of the three) are as shown in Figure 5, the attention level of participant P for each participant will be as shown in Figure 6. For example, the display area of ​​participant 1 is 100 and the total display area is 450, so the attention level of participant 1 in calculation example 1-1 is 100 / 450 = 0.222222. Example 2 Next, a description will be given of Example 2. Example 2 is an example of the short-time attention level estimation unit 120. In a Web conference with many participants, most participants usually mute their microphones, and participants with their microphones on are considered to be attracting attention from the other participants. Therefore, the short-term attention level estimation unit 120 estimates the attention level based on the proportion of time that a participant had their microphone on within a recent specific period of time (e.g., one minute). Examples of calculations performed by the short-time attention level estimation unit 120 include the following calculation examples 2-1 to 2-4. In each of the following formulas, A i is the attention level of participant i, and T short is the most recent specific time span, and st i is the most recent T short is the length of time that participant i in the group has their microphone turned on, and N is the number of participants. In the following calculation, if the denominator is 0 (everyone has their microphones turned off), the attention level of everyone is set to 1 / N. Calculation example 2-1) Calculation example 2-2) Calculation example 2-3) In the above formula, r is a constant greater than or equal to 0. Calculation example 2-4) In the above equation, b is a constant greater than one. Example 2: Specific Example If there are three participants and the values ​​of each variable on the app screen (user terminal 10) of participant P (one of the three) are as shown in Figure 7, participant P's attention level for each participant will be as shown in Figure 8. For example, participant 1's microphone ON time length is 10, and the total time length is 45, so participant 1's attention level in calculation example 2-1 is 10 / 45 = 0.222222. Example 3 Next, a description will be given of Example 3. Example 3 is an example of the long-term attention level estimation unit 130. Even if a person has not spoken recently, if they have a high probability of speaking, they are considered to have a high level of attention. Therefore, the long-term attention level estimation unit 130 estimates the level of attention based on the proportion of time that the microphone was turned on within a long period of time (for example, one hour). Examples of calculations performed by the long-time attention level estimation unit 130 include the following calculation examples 3-1 to 3-4. In each of the following formulas, A i is the attention level of participant i, and T long is a long specific time span, and lt i For example, the most recent long T long is the length of time that participant i in the group has their microphone turned on, and N is the number of participants. In the following calculation, if the denominator is 0 (everyone has their microphones turned off), the attention level of everyone is set to 1 / N. Calculation example 3-1) Calculation example 3-2) Calculation example 3-3) In the above formula, r is a constant greater than or equal to 0. Calculation example 3-4) In the above equation, b is a constant greater than one. Example 3: Specific Example If there are three participants and the values ​​of each variable on the app screen (user terminal 10) of participant P (one of the three) are as shown in Figure 9, participant P's attention level for each participant will be as shown in Figure 10. For example, participant 1's microphone-on time length is 1000, and the total time length is 3300, so participant 1's attention level in calculation example 3-1 is 1000 / 3300 = 0.30303. Example 4 Next, a fourth embodiment will be described. The fourth embodiment is an embodiment regarding the attention level integration unit 140. The attention level integration unit 140 combines at least two of the three attention levels described in the first to third embodiments. When combining two levels, any two levels may be combined. A calculation example in which the three attention levels explained in the first to third embodiments are combined is as follows. A i = α × A1 i + β × A2 i +γ×A3 i In the above formula, A i is the attention level of participant i, and A1 i is the attention level calculated in Example 1, and A2 i is the attention level calculated in Example 2, and A3 i is the attention degree calculated in Example 3. α, β, and γ are weights, at least two of which are not 0, and α+β+γ=1 is satisfied. Example 4: Specific Example It is assumed that the attention levels described in the specific examples of Examples 1 to 3 have been obtained. In other words, it is assumed that the attention levels shown in Figures 6, 8, and 10 have been obtained. It is also assumed that the weights are as shown in Figure 11. Here, the results of calculation examples 1-1, 2-1, and 3-1 are used. In this case, the attention levels of each participant are as shown in Figure 12. For example, for participant 1, 0.1 x 0.222222 + 0.6 x 0.222222 + 0.3 x 0.30303 = 0.2464645. (Example of using attention level) In a Web conference with participants 1, 2, 3, and 4, when estimating the quality experienced by participant 4, a simple average method (i.e., (Σ i EQ i ) / N), it is possible to calculate it by weighting it according to the attention level (i.e., Σ i A i ×EQ i ) enables highly accurate estimation of quality of experience. (Example of hardware configuration) The attention level estimation device 100 described in this embodiment can be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud. That is, the device can be realized by executing a program corresponding to the processing performed by the device using hardware resources such as a CPU and memory built into a computer. The program can be recorded on a computer-readable recording medium (such as a portable memory) and stored or distributed. The program can also be provided via a network such as the Internet or email. Fig. 13 is a diagram showing an example of the hardware configuration of the computer. The computer in Fig. 13 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, all of which are interconnected via a bus B. The computer may further include a GPU. The program that realizes the processing on the computer is provided by a recording medium 1001, such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc. The memory device 1003 reads and stores a program from the auxiliary storage device 1002 when an instruction to start the program is received. The CPU 1004 realizes functions related to the device in accordance with the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, mouse, buttons, a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the results of calculations. (Summary of the embodiment) The attention level estimation device 100 in this embodiment estimates the attention level of each participant in a service in which a plurality of video streams and a plurality of audio streams are integrated and output from each user terminal 10 . The attention level estimation device 100 includes a spatial attention level estimation unit 110 , a short-term attention level estimation unit 120 , a long-term attention level estimation unit 130 , and an attention level integration unit 140 . The spatial attention level estimation unit 110 estimates the spatial attention level based on the display size of each video stream on the integrated screen. The short-term attention level estimation unit 120 estimates the short-term attention level based on state parameters of each integrated video or audio stream. The long-term attention level estimation unit 130 estimates the long-term attention level based on state parameters of each integrated video or audio stream. The attention level integration unit 140 estimates the attention level based on multiple attention levels. (Effects of the embodiment) The technology according to the present embodiment makes it possible to estimate the attention level of each user with high accuracy in a system in which videos of multiple users are displayed on the screens of their respective terminals. It also makes it possible to estimate the user quality of experience taking into account the attention level of each user. <Additional Notes> (Additional note 1) A device used in a system in which a screen on which one or more user images are displayed is displayed on each of a plurality of user terminals, an estimation unit that estimates the degree of attention of the specific user to each user by using either or both of the screen display size of the video of each user on the specific user's screen and the state parameters of each user; An apparatus comprising: (Additional note 2) The estimation unit estimates the attention level based on a screen display size of a video of each user and a screen display size of all users. Item 1. The device according to item 1. (Additional note 3) The estimation unit estimates the attention level using an on / off state of a microphone or a camera of each user as the state parameter. Item 1. The device according to item 1. (Additional note 4) The estimation unit estimates the attention level based on the proportion of time that a microphone or a camera is on in a short period of time or the proportion of time that a microphone or a camera is on in a long period of time. The device described in appended paragraph 3. (Additional note 5) The estimation unit estimates the attention level based on a spatial attention level estimated based on a screen display size of a video of each user and a screen display size of all users, and a temporal attention level estimated using a state parameter of each user. Item 1. The device according to item 1. (Additional note 6) An estimation method executed by a device used in a system in which a screen on which images of one or more users are arranged is displayed on each terminal of a plurality of users, the method comprising: The degree of attention of the specific user to each user is estimated by using either or both of the screen display size of the video of each user on the specific user's screen and the state parameters of each user. Estimation method. (Supplementary Note 7) A non-transitory storage medium storing a program for causing a computer to function as the device according to any one of claims 1 to 4. Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims. 10 User terminal 20 Network 100 Attention level estimation device 110 Spatial attention estimation unit 120 Short-term attention level estimation unit 130 Long-term attention level estimation unit 140 Attention Integration Department 1000 drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device

Claims

1. A device used in a system in which a screen displaying one or more user images is displayed on each of the terminals of multiple users, the device comprising an estimation unit that estimates the level of attention a specific user has for each user by using either or both of the screen display size of each user's image on the specific user's screen and each user's state parameters.

2. The device according to claim 1, wherein the estimation unit estimates the attention level based on the screen display size of the video of each user and the screen display size of all users.

3. The device according to claim 1, wherein the estimation unit estimates the attention level using an on / off state of a microphone or camera of each user as the state parameter.

4. The device according to claim 3, wherein the estimation unit estimates the attention level based on the proportion of time that a microphone or camera is on in a short period of time, or the proportion of time that a microphone or camera is on in a long period of time.

5. The device described in claim 1, wherein the estimation unit estimates the attention level based on a spatial attention level estimated based on the screen display size of each user's video and the screen display size of all users, and a temporal attention level estimated using state parameters of each user.

6. An estimation method executed by a device used in a system in which a screen displaying one or more user images is displayed on each terminal of a plurality of users, the method estimating the level of attention a specific user has for each user by using either or both of the screen display size of each user's image on the specific user's screen and each user's state parameters.

7. A program for causing a computer to function as the device according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Communication terminal and display method thereof

    JP2007150919A

  • Information providing system, information providing method, information display device, and information display method

    JP2011233095A

  • Communication control apparatus, conference system and program

    JP2017152952A