Information processing device, information processing method, and program

JP7927499B2Active Publication Date: 2026-10-01CANON KK
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2022126243
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2026-10-01
Estimated Expiration
2042-08-08

AI Technical Summary

Benefits of technology

【0008】 本開示により、コンテンツデータを再生した際、視聴するユーザのリアクションに応じて、コンテンツデータの再生を制御することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927499000001
    Figure 0007927499000001
  • Figure 0007927499000002
    Figure 0007927499000002
  • Figure 0007927499000003
    Figure 0007927499000003
Patent Text Reader

Abstract

To control reproduction of content data according to the reaction of a viewing user when reproducing the content data.SOLUTION: A content data acquisition unit acquires a lecture image captured by a first imaging device and stores the lecture image in a main memory. A first object extraction unit extracts a prescribed object in the acquired lecture image and stores the object in the main memory. A reaction data acquisition unit acquires a hand side image captured by a second imaging device via an external connection I / F and stores the hand side image in the main memory. A second object extraction unit extracts a prescribed object from the acquired hand side image and stores the object in the main memory as second object information. A grasping degree determination unit determines a grasping degree on the basis of the extracted first object information and second object information and stores information about the grasping degree in the main memory. A reproduction control unit controls display of the lecture image in a display unit on the basis of the grasping degree.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing technology for performing playback control of content data.

Background Art

[0002] In recent years, with the development of Internet technology, the introduction of e-learning systems via real-time distribution and on-demand distribution of lectures has been progressing in educational settings. On the other hand, in many e-learning systems, the playback method for content data such as lectures is monotonous, and information is only transmitted in one direction, which makes the learning attitude of students passive and tends to reduce learning efficiency.

[0003] Patent Document 1 discloses a technology that associates the playback of lecture videos with the recording status taken by students based on the finding that learning efficiency can be improved by assigning students the active act of taking notes in a notebook or the like. Specifically, the playback speed of the video is controlled according to the key input speed when the student takes notes via a keyboard.

Prior Art Literature

Patent Literature

[0004]

Patent Document 1

Summary of the Invention

Problem to be Solved by the Invention

[0005] However, in Patent Document 1, the playback speed of the lecture video is controlled according to the key input speed, but the content of the key input does not necessarily match the playback content of the lecture video, so there is a problem that it is difficult for students to focus on the lecture video while taking notes.

[0006] Therefore, this disclosure aims to control the playback of content data in response to the user's reaction when the content data is played. [Means for solving the problem]

[0007] This disclosure relates to an information processing apparatus comprising: a first acquisition means for acquiring content data including at least one of image data and audio data; a control means for controlling the playback of the content data to be presented by a presentation device including at least one of a display means and an audio output means; and the content data presented by the presentation device. For the aforementioned content data The system comprises a second acquisition means for acquiring reaction data including at least one of image data and audio data recording a user's reaction, and an extraction means for extracting object information relating to at least one object of characters, symbols, and figures from the content data and the reaction data, respectively, wherein the control means controls the playback of the content data based on the object information of the content data and the object information of the reaction data obtained from the extraction means. [Effects of the Invention]

[0008] This disclosure makes it possible to control the playback of content data in response to the user's reaction when the content data is played. [Brief explanation of the drawing]

[0009] [Figure 1] A diagram showing an example system configuration including user terminals according to the first and second embodiments. [Figure 2] Block diagrams showing examples of hardware configurations of user terminals according to the first and second embodiments. [Figure 3] A block diagram showing an example of the functional configuration of a user terminal in the first embodiment. [Figure 4] A flowchart of the process performed by the user terminal in the first embodiment. [Figure 5]A conceptual diagram illustrating an example of lecture images and text information within a lecture. [Figure 6] A conceptual diagram illustrating an example of a handheld image and handwritten text information. [Figure 7] A flowchart of the processing performed by the regeneration control unit in the first embodiment. [Figure 8] A conceptual diagram illustrating an example of a lecture image whose display is controlled by the playback control unit in the first embodiment. [Figure 9] A flowchart of the processing performed by the regeneration control unit in the second embodiment. [Figure 10] A conceptual diagram illustrating an example of a predetermined instruction in a handheld image. [Figure 11] A diagram showing an example system configuration including a user terminal according to the third and fourth embodiments. [Figure 12] Block diagrams showing examples of hardware configurations for user terminals according to the third and fourth embodiments. [Figure 13] A block diagram showing an example of the functional configuration of a user terminal in the third and fourth embodiments. [Figure 14] A flowchart of the processing performed by the user terminal in the third embodiment. [Figure 15] A flowchart of the processing performed by the regeneration control unit in the third embodiment. [Figure 16] A conceptual diagram illustrating an example of gradual emphasis in the third embodiment. [Figure 17] A conceptual diagram illustrating an example of the types of emphasis in the third and fourth embodiments. [Figure 18] A flowchart of the processing performed by the user terminal in the fourth embodiment. [Modes for carrying out the invention]

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. The following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be combined arbitrarily. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and overlapping descriptions are omitted.

[0011] [First Embodiment] A configuration example of the e-learning system according to the present embodiment will be described with reference to FIG. 1. In the system of the present embodiment, in a classroom where a blackboard 106 is installed, a lecture given by a lecturer 107 using the blackboard 106 is photographed using a first imaging device 101. The first imaging device 101 is connected to an information processing device 102 via a network. Then, the information processing device 102 is connected to a user terminal 103 via another network. An image captured by the first imaging device 101 is transmitted to the information processing device 102 for image analysis and image editing etc. After being subjected to the image processing described above, the image is transmitted to the user terminal 103. The information processing device 102 is a computer device on a network, such as a cloud service or a server for internal use. The user terminal 103 is an information processing device such as a PC (personal computer), a smartphone, or a tablet, and has a display device for displaying the transmitted image.

[0012] Furthermore, the user terminal 103 is connected to a second imaging device 104, and the second imaging device 104 captures an image of a notebook 105 at hand of a student 108. While viewing the image displayed on the user terminal 103, the student 108 actively learns by copying the content of the blackboard 106 into the notebook 105, writing down important points, and the like. The notebook 105 may be a commonly used learning notebook, a print sheet distributed in advance by the lecturer, or the like.

[0013] Next, an example of the hardware configuration of the user terminal 103 according to this embodiment will be described using the block diagram in Figure 2. The hardware of the user terminal 103 consists of a CPU 201, main memory 202, storage unit 203, operation unit 204, display unit 205, communication unit 206, external connection I / F 207, and bus 208.

[0014] The CPU 201 executes various processes using computer programs and data stored in the main memory 202. In doing so, the CPU 201 controls the overall operation of the user terminal 103 and executes or controls the various processes described as being performed by the user terminal 103.

[0015] The main memory 202 has an area for storing computer programs and data loaded from the storage unit 203, data received from the outside by the communication unit 206, and so on. Furthermore, the main memory 202 has a work area used by the CPU 201 when executing various processes. In this way, the main memory 202 can provide various areas as appropriate.

[0016] The memory unit 203 stores the OS (operating system) and computer programs and data that cause the CPU 201 to execute or control various processes described as being performed by the user terminal 103. The computer programs and data stored in the memory unit 203 are loaded into the main memory 202 as appropriate according to the control of the CPU 201 and become subject to processing by the CPU 201. Non-volatile memory such as a silicon disk can be applied to the memory unit 203.

[0017] The control unit 204 is a user interface that includes a keyboard, mouse, buttons, mode dial, switches, levers, and a touch panel screen, and allows the trainee 108 to input various instructions to the CPU 201 by operating it.

[0018] The display unit 205 has a screen such as an LCD screen or a touch panel screen, and displays images including characters and diagrams processed by the CPU 201. If the display unit 205 has a touch panel screen, the operation input entered by the trainee 108 by operating the touch panel screen is notified to the CPU 201.

[0019] The communication unit 206 is a device compliant with communication standards such as local area networks and IEEE 802.11, and can perform data communication with external devices.

[0020] The External Connection I / F207 is a device compliant with interface standards such as USB, and it enables data communication with peripheral devices such as imaging devices.

[0021] Furthermore, the CPU 201, main memory 202, storage unit 203, operation unit 204, display unit 205, communication unit 206, and external connection I / F 207 are all connected to the bus 208.

[0022] Furthermore, the user terminal 103 may also be equipped with a speaker, and may output audio processed by the CPU 201.

[0023] Figure 3 shows a block diagram illustrating an example of the functional configuration of the user terminal 103 according to this embodiment. The user terminal 103 includes a content data acquisition unit 301 that acquires content data including at least one of image data such as lecture images and audio data obtained by capturing images with the first imaging device 101. The user terminal 103 also includes a first object extraction unit 302 that extracts information about predetermined objects such as characters, symbols, and figures (hereinafter referred to as first object information) from the lecture images. The user terminal 103 also includes a reaction data acquisition unit 303 that acquires hand images obtained by capturing images with the second imaging device 104 as reaction data. The user terminal 103 also includes a second object extraction unit 304 that extracts information about predetermined objects such as characters, symbols, and figures (hereinafter referred to as second object information) from the hand images. The user terminal 103 also includes a comprehension determination unit 305 that determines the degree of comprehension of the content data playback based on the student's writing status, based on the extraction results of the first object extraction unit 302 and the second object extraction unit 304. The user terminal 103 also includes a playback control unit 306 that controls the method of playing back lecture images according to the user's level of understanding.

[0024] Figure 4 shows a flowchart of the processing performed by the user terminal 103 according to this embodiment. In the following description, the functional unit shown in Figure 3 will be described as the main component of the processing. However, in reality, the CPU 201 executes a computer program that enables the CPU 201 to implement the functions of the functional unit of the user terminal 103, thereby realizing the functions of the functional unit of the user terminal 103. Note that the functional unit shown in Figure 3 may also be implemented in hardware.

[0025] In S401, the content data acquisition unit 301 acquires lecture images captured by the first imaging device 101 via the communication unit 206 and stores them in the main memory 202. The lecture images to be stored may be still images or a series of frames of a moving image.

[0026] In S402, the first object extraction unit 302 extracts predetermined objects from the lecture image acquired in S401 and stores them in the main memory 202. Here, predetermined objects include handwritten characters and diagrams written by the lecturer 107 on a blackboard or whiteboard, time information for each object's extraction, and coordinate information for each object. Optical Character Recognition (OCR) technology is a widely known method for extracting characters and diagrams from images. OCR technology performs pre-processing such as noise reduction and tilt correction on the image to be processed, then performs layout analysis such as columns and tables, separates character areas, and finally extracts characters through pattern recognition and feature detection. It is also possible to improve the accuracy of character extraction by using methods such as background subtraction, which extracts the difference area by comparing a reference background image with the input image. In background subtraction, the reference background image can be generated by calculating the time average for each pixel from temporally consecutive captured images, or by taking a picture of an image without a subject in advance. This makes it possible to extract not only characters written on blackboards, etc., but also non-character objects such as symbols and diagrams. Furthermore, this method can extract not only handwritten characters and diagrams, but also printed characters, symbols, and diagrams, as well as computer-generated characters and diagrams. Alternatively, audio data may be acquired instead of image data such as lecture images.

[0027] Figure 5 shows an example of first object information related to the extraction results of a lecture image. Figure 5(a) is a diagram showing an example of a lecture image acquired in S401. The lecture image shows the instructor 107 writing the strings "ABCDE" and "XYZ" 501 on the blackboard 106. Note that the strings 501 shown here are for illustrative purposes only and may be more complex strings, phrases, symbols, or figures that can be extracted using the method described above. Figure 5(b) is a table showing the first object information extracted from the lecture image shown in Figure 5(a). The table shown in Figure 5(b) stores information about the extracted characters, the time code (time information) when each character was first extracted, and the coordinates of the extracted characters in the lecture image. Here, it shows when (time code) and where (coordinates) in the lecture image each character in the strings "ABCDE" and "XYZ" 501 was extracted. This first object information is stored in main memory 202.

[0028] In S403, the reaction data acquisition unit 303 acquires the hand image recorded by the second imaging device 104, which is received via the external connection I / F 207, and stores it in the main memory 202. Alternatively, instead of the hand image, audio data may be acquired in which the student 108 reads aloud the text information contained in the lecture image. Typical examples of using audio data include language pronunciation practice.

[0029] In S404, the second object extraction unit 304 extracts predetermined objects from the hand image acquired in S403 and stores them in the main memory 202 as second object information. Here, the second object information consists of objects such as characters, symbols, and diagrams written by the student 108 in their notebook, and a time code indicating the time when each object was extracted. The extraction of the second object information can be done in the same way as in S402, so the explanation is omitted.

[0030] On the other hand, let's consider the case where the reaction data, which contains user reactions acquired in S401 and S403, is audio data. Speech recognition technology is a widely known method for extracting text from audio data. Speech recognition technology removes noise from the waveform components of the audio data, extracts the phonemes (the smallest components of sound) by pattern recognition while cutting out the waveform, and then extracts text by matching which words the sound is closest to.

[0031] Figure 6 shows the results of extracting the local image. do An example of second object information is shown. Figure 6(a) is an example of a handheld image acquired in S403, showing the string "ABC" 601 written by student 108 on notebook 105. Figure 6(b) is a table showing the second object information extracted from the handheld image shown in Figure 6(a). The table in Figure 6(b) stores the extracted character information and the time code of when it was first extracted. Here, it shows when (time code) and from where (coordinates) in the handheld image each character in the string "ABC" 601 was extracted. This second object information is stored in main memory 202.

[0032] In S405, the comprehension determination unit 305 determines the comprehension level based on the extracted first object information and second object information, and stores the information regarding the comprehension level in the main memory 202. Comprehension level is an index that evaluates how well the student 108 understands the content explained by the instructor 107. Specifically, it is an index determined based on the timing of object presentation by the instructor 107, the timing of writing by the student 108, and the amount of objects that the student 108 should write down. In this embodiment, the comprehension level is determined based on the extraction time of the object extracted from the content data corresponding to the latest object extracted from the reaction data, and the amount of objects extracted from the content data after that extraction time. For example, if there is a long time between the instructor 107 writing on the board and the student 108 writing down the content, or if there is a large amount of text that the student 108 should write down, the comprehension level can be determined to be low. The comprehension level may be determined using rules based on parameters such as thresholds, or it may be determined using inference based on machine learning.

[0033] Here, we will explain how to determine the degree of understanding based on the example of first object information shown in Figure 5 and the example of second object information shown in Figure 6. According to Figure 6(b), the last character extracted in the second object information is "C", and the time code for the time of extraction is 00:00:32:00. Also, according to Figure 5(b), the time code for the time when the character "C" was extracted in the first object information is 00:00:03:00. As a result, the time difference between when instructor 107 writes the character "C" on the blackboard 106 and when student 108 writes the character "C" in their notebook is 00:00:29:00. Furthermore, at the time when the last character written by student 108 is extracted, the remaining characters that student 108 should write are "DEXYZ", which amounts to 5 characters.

[0034] In S406, the playback control unit 306 controls the display of the lecture image on the display unit 205 based on the level of understanding determined in S405. In this embodiment, the display area of ​​the lecture image on the display unit 205 is controlled based on the level of understanding.

[0035] Figure 7 shows a flowchart illustrating the process of controlling the display area of ​​the lecture image based on the level of understanding according to this embodiment.

[0036] In S701, the playback control unit 306 acquires the degree of understanding determined in S405.

[0037] In S702, the regeneration control unit 306 determines whether the level of understanding is lower than a predetermined reference value. If the level of understanding is less than the reference value, the process proceeds to S703; if the level of understanding is equal to or greater than the reference value, the process proceeds to S704.

[0038] In S703, the playback control unit 306 determines whether a predetermined time has elapsed since the state of low comprehension. If the predetermined time has elapsed, the process proceeds to S706; otherwise, it proceeds to S705.

[0039] In S704, the playback control unit 306 determines a display area centered on the image area containing the latest object, which is the character extracted as the first object information, and controls the display unit 205 to enlarge and display the determined display area from the lecture image acquired in S401.

[0040] In S705, the playback control unit 306 determines a display area centered on the image area containing the most recent character extracted as second object information, and controls the display unit 205 to enlarge and display the determined display area from the lecture image acquired in S401.

[0041] In S706, the playback control unit 306 determines a display area centered on the image area containing the character corresponding to the most recent object extracted as second object information. Then, in S706, it controls the display unit 205 to enlarge the determined display area from the lecture image acquired in S401 and to superimpose warning information indicating a low level of comprehension onto the display unit 205.

[0042] Figure 8 is a conceptual diagram illustrating an example of a lecture image whose display is controlled by the playback control unit 306. Figure 8(a) is an example of a display centered on the image region containing the most recently extracted character as first object information. According to Figure 5(b), the most recently extracted character as first object information is "Z", and the display is controlled so that the image region containing "Z" is located near the center. Figure 8(b) is an example of a display centered on the image region containing the most recently extracted character as second object information. According to Figure 6(b), the most recently extracted character as second object information is "C", and the display is controlled so that the image region containing "C" is at the center. Figure 8(c) is an example of displaying warning information indicating a low level of comprehension. In the example display, "[Caution] You are behind on taking notes!" is displayed at the bottom of the image.

[0043] Thus, according to this embodiment, the display area of ​​the lecture image can be controlled with high precision based on the level of understanding of the learners 108, thereby making it easier for the learners 108 to concentrate on the next area of ​​interest and improving the learning effect.

[0044] [Second Embodiment] In the following embodiments, including this embodiment, the differences from the first embodiment will be described, and unless otherwise specified below, they will be the same as the first embodiment. In the first embodiment, a method for controlling the display area of ​​an image based on the comprehension level of the students 108 was described. In this embodiment, a method for controlling the playback speed of a lecture image based on the comprehension level of the students 108 will be described.

[0045] The examples of the functional configurations of the user terminal 103 are the same as in the first embodiment and are therefore omitted. Also, the lecture image acquired in S401 may be audio data instead of image data. In that case, the first object information extracted in S402 is the characters extracted from the audio data and the time information related to those characters.

[0046] Here, we will explain in detail the process of S406, which controls the playback speed of the lecture image on the display unit 205 based on the level of comprehension, following the flowchart in Figure 9.

[0047] In S701, the playback control unit 306 acquires the degree of understanding determined in S405.

[0048] In S702, the regeneration control unit 306 determines whether the level of understanding is lower than a predetermined reference value. If the level of understanding is less than the reference value, the process proceeds to S902; if the level of understanding is equal to or greater than the reference value, the process proceeds to S901.

[0049] In S901, the playback control unit 306 controls the playback speed of the lecture image acquired in S401 to slow it down. For example, the playback speed may be controlled to slow down in proportion to the length of time between when the instructor 107 writes on the board and when the student 108 takes notes of the content. Alternatively, if the data acquired in S401 and S403 is audio data, speech analysis technology may be used to extract the tempo and rhythm of the speech corresponding to the characters, and the playback speed may be controlled so that it matches the tempo and rhythm of the student 108's note-taking.

[0050] In S902, the playback control unit 306 determines whether or not there is instruction information to increase the playback speed in the local image acquired in S403. If there is predetermined instruction information, the process proceeds to S903; otherwise, it proceeds to S904.

[0051] In S903, the playback control unit 306 determines the display speed to increase the playback speed and controls the display of the lecture image acquired in S401. For example, the playback speed may be controlled to increase until the amount of character information that the student 108 needs to write down exceeds a predetermined value.

[0052] In S904, the playback control unit 306 determines the display speed so that the playback speed is the normal speed, and controls the display of the lecture image acquired in S401.

[0053] Figure 10 is a conceptual diagram showing an example of predetermined instruction information provided by participant 108. In Figure 10, the most recent character extracted as the second object information of participant 108 is "Z", indicating that the level of understanding is not slow. In the lower left of the screen, participant 108's left hand 1002 is displayed, showing a pose with the thumb and index finger extended. In this embodiment, an image of a left hand with the thumb and index finger extended is pre-registered as predetermined instruction information in the main memory 202 or storage unit 203, and the predetermined instruction information can be detected by the background subtraction method or pattern recognition described above.

[0054] Thus, according to this embodiment, the playback speed of the lecture images can be controlled based on the level of understanding of the learners 108, thereby enabling the lecture to progress with greater precision according to the learning pace of the learners 108, and improving the learning effect.

[0055] [Third Embodiment] The objective of this embodiment is to provide a user terminal that enhances the learning effectiveness of learners 108 with visual impairments. It is known that 2.4% of students in regular classrooms in Japan have significant difficulties with reading or writing. One of the factors that makes reading and writing difficult is problems with visual functions such as fixation, tracking, fixation point shifts, and visual search. When training visual functions, it is desirable to impose training of an appropriate intensity. When training learning and visual functions by taking notes from the blackboard during regular lessons, it can be said that learners 108 should be assisted so that they do not have too much trouble identifying the parts written on the blackboard. However, in the first embodiment, if the level of comprehension is low, the image area corresponding to the latest character extracted as second object information is always enlarged, which may be excessive assistance and may reduce the training effect. Therefore, in this embodiment, by appropriately indicating the parts written on the blackboard that should be written in the notebook according to the learner 108's condition, it is possible to impose training suitable for learner 108 and enhance the training effect of learner 108.

[0056] Figure 11 shows an example configuration according to this embodiment. In the system of this embodiment, in addition to the first embodiment, the user terminal 103 places a third imaging device 1101 capable of capturing the face of the student 108 next to the screen.

[0057] Figure 12 shows an example of the hardware configuration of the user terminal 103 according to this embodiment. In the system of this embodiment, in addition to the second imaging device 104 of the first embodiment, data communication with the third imaging device 1101 is performed via the external connection I / F 207.

[0058] Figure 13 shows an example of the functional configuration of the user terminal 103. In addition to the configuration of the first embodiment, the system of this embodiment further includes a face image acquisition unit 1301 and a gaze point information extraction unit 1302. Details of the functions of the face image acquisition unit 1301 and the gaze point information extraction unit 1302 will be described later.

[0059] Figure 14 shows a flowchart illustrating the processing performed by the user terminal 103 according to this embodiment.

[0060] In S1401, the facial image acquisition unit 1301 acquires facial image data of the trainee 108 from the third imaging device 1101 received via the external connection I / F 207 and stores it in the main memory 202.

[0061] In S1402, the gaze point information extraction Unit 1302 applies a gaze estimation method using deep learning to the facial image data of participant 108 acquired in S1401, thereby deriving the position information of participant 108's gaze point from changes in participant 108's posture and eye movements. The gaze point information extraction unit 1302 stores the gaze point information, including the extracted gaze point position information of participant 108, in the main memory 202 in chronological order.

[0062] In S1403, the playback control unit 306 controls the display to highlight a portion of the lecture image based on the level of understanding and the point of focus information.

[0063] Figure 15 shows a flowchart illustrating the process of controlling the display to highlight a portion of the lecture image based on the level of understanding and gaze point information according to this embodiment.

[0064] In S1501, the playback control unit 306 uses the gaze point information to determine whether the gaze point of the participant 108 is on the screen of the display unit 205 of the user terminal 103, and whether the amount of movement of the gaze point per unit time is greater than or equal to a predetermined threshold, and whether the participant 108's gaze is not fixed. not present If this is the case, proceed to S1502; if the participant 108's gaze point is not on the screen of the display unit 205, or if their gaze is fixed, proceed to S1505.

[0065] In S1502, the playback control unit 306 determines whether the lecture image acquired in S401 is being highlighted character by character. Details of character-level highlighting will be described later. If image processing for character-level highlighting has not been applied, the process proceeds to S1503; if image processing for character-level highlighting has been applied, the process proceeds to S1504.

[0066] In S1503, the playback control unit 306 displays the lecture image acquired in S401 as is.

[0067] In S1504, the playback control unit 306 gradually highlights the area of ​​the lecture image acquired in S401 that should be written to in the notebook 105, so that the student 108 does not have too much trouble identifying it. The playback control unit 306 highlights the area containing the latest object, which is the character, extracted as the second object information, either line by line or character by character. The playback control unit 306 first highlights line by line. If the student 108's focus remains unsettled for a certain period of time, it moves to highlighting character by character. If the student 108's focus remains unsettled for a further period of time, the display area centered on the image area containing the latest character extracted as the second object information, as shown in the first embodiment, is enlarged and then highlighted. When the student 108 finishes writing down a block of text, such as the strings "ABCDE" and "XYZ", line by line, the highlighting ends.

[0068] Furthermore, the playback control unit 306 can be pre-configured by the learner 108 to always perform predetermined highlighting on a line-by-line or character-by-character basis, and can be configured not to perform gradual highlighting control. In addition, the playback control unit 306 may dynamically change the time it takes to transition to character-by-character highlighting based on indicators that evaluate the readability of the characters, such as the font size of the characters written on the blackboard.

[0069] Figure 16 shows a conceptual diagram illustrating an example of stepwise highlighting according to this embodiment. Figure 16(a) is the lecture image acquired in S401 and displayed in S1503, showing the entire blackboard 106. Figure 16(b) is the lecture image highlighted line by line and displayed in S1504, with image processing performed to increase the brightness of line 1601 containing the latest character extracted as second object information. Figure 16(c) is the lecture image highlighted character by character and displayed in S1504, with image processing performed to increase the brightness of characters 1602 from the latest character extracted as second object information within line 1601. Figure 16(d) is the lecture image highlighted character by character and displayed in S1504, with an enlarged display area centered on the image area containing the latest character extracted as second object information, in addition to the highlighting shown in Figure 16(c).

[0070] Figure 17 is a conceptual diagram showing an example of highlighting types, and the playback control unit 306 can select and use a suitable highlighting from among multiple highlighting types.

[0071] In the highlighting shown in Figure 17(a), the brightness of line 1601 containing the most recently extracted character as second object information is increased for the blackboard 106. Highlighting may also be applied on a character-by-character basis.

[0072] In the highlighting shown in Figure 17(b), the brightness and color of line 1601 containing the most recent character extracted as second object information are changed on the blackboard 106. Since it is known that visibility can be improved by using a colored filter depending on the visual function problem, the color is also changed. The color to be changed is set in advance by the student 108. Highlighting may also be done on a character-by-character basis.

[0073] In the highlighting shown in Figure 17(c), the brightness and color of line 1601 containing the most recently extracted character as second object information are changed on the blackboard 106, and a predetermined width is filled in so that the characters above and below it become invisible. This is a reproduction of a reading ruler that may be used by people with visual impairments. The color to be changed is set in advance by the student 108. To highlight on a character-by-character basis, for example, a reading ruler adjusted to the length of the character can be slid to the relevant location.

[0074] In the highlighting shown in Figure 17(d), a marker 1701 is drawn at the beginning of line 1601 containing the most recently extracted character as second object information on the blackboard 106. The position where the marker 1701 is drawn is not limited to the beginning of the line; the marker 1701 may be drawn within the line, or the marker 1701 may be used to highlight characters individually.

[0075] In the highlighting shown in Figure 17(e), the characters included in the second object information are removed from the blackboard 106. Using the first object information extracted in S402 and the second object information extracted in S404, the characters included in the second object information are extracted from the characters included in the first object information, and the area indicated by the coordinates of each extracted character is filled with the color surrounding that character. However, if the character is hidden by the instructor 107 or other objects, it does not need to be filled in.

[0076] Thus, according to this embodiment, it is possible to control the display of lecture images to highlight them based on the learner's level of comprehension and gaze point information, thereby improving the training effect for learners 108 with visual impairments.

[0077] Similar to the second embodiment, this embodiment may also detect predetermined instructions from the hand poses of the trainees 108 and control the display / hide, enlargement / reduction, movement, rotation, etc., of the highlighted area based on the trainees 108's instructions.

[0078] [Fourth Embodiment] The objective of this embodiment, like that of the third embodiment, is to provide a user terminal that enhances the training effect for learners 108 with visual impairments. It is desirable that learners 108 with visual impairments be able to quickly find the parts of the blackboard that instructor 107 is explaining, even if they do not take notes from the blackboard. However, if the parts of the blackboard that instructor 107 is explaining are constantly highlighted, it will not only lead to a decrease in training effectiveness, as described in the third embodiment, but there is also a concern that the transitions in highlighting will be noticeable and disrupt the learners 108's concentration. Therefore, in this embodiment, similar to the third embodiment, the training effect for learners 108 is enhanced by appropriately indicating the parts of the blackboard that instructor 107 is explaining, according to the learners 108's condition.

[0079] The system configuration example and hardware configuration example according to this embodiment are the configurations shown in Figures 11 and 12 with the second imaging device 104 removed. Furthermore, the functional configuration example according to this embodiment is the configuration shown in Figure 13 with the reaction data acquisition unit 303 and the second object extraction unit 304 removed.

[0080] Figure 18 is a flowchart showing the processing performed by a user terminal 103 having the example functional configuration shown in Figure 3.

[0081] In S1801, the content data acquisition unit 301 acquires the lecture image received by the communication unit 206 and stores it in the main memory 202. Furthermore, the content data acquisition unit 301 acquires the audio data received by the communication unit 206 and stores it in the main memory 202.

[0082] In S1802, the first object extraction unit 302 extracts predetermined objects from the lecture image acquired in S1801, similar to the first embodiment, and stores them in the main memory 202 as first object information. Furthermore, the first object extraction unit 302 recognizes words spoken by the lecturer 107 from the audio data acquired in S1801 using the same speech recognition technology as in the first embodiment, and compares the characters constituting the recognized words with the characters in the first object information. The first object extraction unit 302 then extracts character information corresponding to each character constituting the words spoken in the explanation from within the first object information. Alternatively, the first object extraction unit 302 may analyze the lecture image to extract gestures of the lecturer 107 and extract character information having coordinate information near the position indicated by the extracted gestures.

[0083] In S1803, the regeneration control unit 306 is configured similarly to S1501 in the third embodiment. Note Determine if the viewpoint is fixed on the screen of the user terminal 103. If the point of focus is fixed on the screen of the user terminal 103, proceed to S1503; otherwise, proceed to S1804.

[0084] In S1804, the playback control unit 306 performs highlighting, similar to S1504 in the third embodiment. It controls the image data acquired in S1801 to highlight the words or lines containing the words that were matched in S1802.

[0085] Thus, according to this embodiment, the part of the whiteboard that the instructor 107 is explaining can be shown only when the student 108's gaze is not fixed, thereby providing a user terminal that enhances the training effect for students 108 with visual impairments.

[0086] The numerical values, processing timing, processing order, display screen configuration, displayed information, display format, and the entity responsible for processing used in the above explanation are given as examples for the purpose of providing a concrete explanation, and are not intended to be limited to these examples alone.

[0087] Furthermore, you may use some or all of the embodiments described above in appropriate combinations, or you may use some or all of the embodiments selectively.

[0088] (Other embodiments) In any of the configurations 1 to 4 of this embodiment, the user terminal 103 performed all processing, but some of the processing of the present invention may be performed by the information processing device 102. For example, by transmitting image data acquired from the second imaging device 104 to the information processing device 102, the information processing device 102 may extract the second object information and determine the degree of recognition. Furthermore, the display area of ​​the image may be controlled or the playback speed may be controlled based on the degree of recognition, and the controlled image data may be transmitted to the user terminal 103.

[0089] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0090] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention.

[0091] This disclosure includes the following configurations and methods:

[0092] (Composition 1) An information processing device comprising: a first acquisition means for acquiring content data including at least one of image data and audio data; a control means for controlling the playback of the content data to be presented by a presentation device; a second acquisition means for acquiring reaction data including at least one of image data or audio data recording the user's reaction presented by the presentation device; and an extraction means for extracting object information relating to at least one object of characters, symbols, and figures from the content data and the reaction data, respectively, wherein the control means controls the playback of the content data based on the object information of the content data and the object information of the reaction data obtained from the extraction means.

[0093] (Configuration 2) The information processing apparatus according to Configuration 1, further comprising a determination means for determining the user's understanding of the content data playback content based on object information of the content data obtained from the extraction means and object information of the reaction data, wherein the control means controls the playback of the content data based on the understanding level.

[0094] (Composition 3) The information processing apparatus according to configuration 2, characterized in that the extraction means acquires the extraction time of each object information, and the determination means determines the degree of understanding based on the difference between the extraction time of an object extracted from the content data and the extraction time of an object extracted from the reaction data corresponding to the object extracted from the content data.

[0095] (Composition 4) The information processing device according to configuration 3, characterized in that it determines the degree of understanding based on the extraction time of an object extracted from the content data corresponding to the latest object extracted from the reaction data, and the amount of objects extracted from the content data after the extraction time.

[0096] (Composition 5) The information processing apparatus according to any one of configurations 2 to 4, characterized in that the extraction means acquires positional information in the image data of an object extracted from the content data when the content data includes the image data, and the playback control, based on the degree of recognition and the positional information, enlarges and displays the image region corresponding to the latest object extracted from the reaction data on the display device when playing back the image data.

[0097] (Composition 6) The information processing apparatus according to any one of configurations 2 to 4, characterized in that the extraction means acquires positional information in the image data of an object extracted from the content data when the content data includes the image data, and the playback control, when the level of understanding is lower than a threshold, displays an image region corresponding to the latest object extracted from the reaction data near the center of the display screen of the presentation device when playing back the image data.

[0098] (Composition 7) The information processing device according to configuration 5 or 6, characterized in that, when the recognition level remains below a threshold for a predetermined period of time, the playback control superimposes warning information when playing back the image data.

[0099] (Composition 8) The information processing apparatus according to any one of configurations 5 to 7, characterized in that the extraction means further extracts predetermined instruction information from the reaction data, and the playback control, when the extraction means has extracted the instruction information, highlights the latest object extracted from the reaction data of the content data when playing back the image data.

[0100] (Composition 9) An information processing apparatus according to any one of configurations 5 to 8, characterized in that: a third acquisition means for acquiring the user's face image data; if the content data includes image data, a derivation means for deriving position information of the point of focus on the display screen of the presentation device that displays the image data that the user is looking at, based on the face image data; and the playback control, when the amount of movement of the point of focus per unit time is greater than or equal to a threshold, highlights the latest object extracted from the reaction data when playing back the image data.

[0101] (Composition 10) The information processing apparatus according to configuration 9, characterized in that the playback control highlights a part of the image data by changing the brightness of the object, changing the color of the object, displaying a leading ruler on the object, displaying a marker on the object, or deleting an object extracted from the reaction data.

[0102] (Composition 11) The information processing device according to any one of configurations 2 to 10, characterized in that the playback control changes the playback speed of the content data based on the degree of understanding.

[0103] (Composition 12) The information processing apparatus according to any one of configurations 1 to 11, characterized in that the extraction means extracts predetermined instruction information from the reaction data, and the playback control changes the playback speed of the content data based on the instruction information.

[0104] (Composition 13) An information processing method comprising: a first acquisition step of acquiring content data including at least one of image data and audio data; a control step of controlling the playback of the content data to be presented by a presentation device; a second acquisition step of acquiring reaction data including at least one of image data or audio data recording the user's reaction presented by the presentation device; and an extraction step of extracting at least one piece of information from characters, symbols, and figures from the content data and the reaction data, respectively, wherein the control step controls the playback of the content data based on the object information of the content data and the object information of the reaction data obtained in the extraction step.

[0105] (Composition 14) A program for causing a computer to function as an information processing device as described in any one of items 1 to 12. [Explanation of Symbols]

[0106] 103 User terminal 301 Content Data Acquisition Unit 302 First Object Information Extraction Unit 303 Reaction Data Acquisition Unit 304 Second Object Information Extraction Unit 305 Grasping degree determination section 306 Regeneration Control Unit

Claims

1. A first acquisition means for acquiring content data including at least one of image data and audio data, Control means for controlling the playback of the content data to be presented by a presentation device including at least one of a display means and an audio output means, A second acquisition means for acquiring reaction data, which includes at least one of image data and audio data recording the user's reaction to the content data presented by the presentation device, Extraction means for extracting object information relating to at least one object of characters, symbols, and figures from the content data and the reaction data, respectively. Equipped with, The control means controls the playback of the content data based on the object information of the content data obtained from the extraction means and the object information of the reaction data. An information processing device characterized by the following:

2. The system further comprises a determination means for determining the user's understanding of the content data's playback content based on the object information of the content data obtained from the extraction means and the object information of the reaction data, The control means controls the playback of the content data based on the degree of understanding. The information processing apparatus according to feature 1.

3. The extraction means obtains the extraction time of each object information, The determination means determines the degree of understanding based on the difference between the extraction time of an object extracted from the content data and the extraction time of an object extracted from the reaction data corresponding to the object extracted from the content data. The information processing apparatus according to feature 2.

4. The degree of understanding is determined based on the extraction time of the object extracted from the content data corresponding to the latest object extracted from the reaction data, and the amount of objects extracted from the content data after that extraction time. The information processing apparatus according to claim 3.

5. If the content data includes the image data and the presentation device includes the display means, the extraction means acquires the position information in the image data of the object extracted from the content data. Based on the degree of understanding and the position information, the control means, when reproducing the image data, causes the display device to enlarge and display the image region corresponding to the latest object extracted from the reaction data. The information processing apparatus according to feature 2.

6. If the content data includes the image data and the presentation device includes the display means, the extraction means acquires the position information in the image data of the object extracted from the content data. When the recognition level is lower than a threshold, the control means causes the image region corresponding to the latest object extracted from the reaction data to be displayed near the center of the display screen of the display means when playing back the image data. The information processing apparatus according to feature 2.

7. The control means, when the level of understanding remains below a threshold for a predetermined period of time, superimposes warning information onto the image data during playback. The information processing apparatus according to feature 5.

8. The extraction means further extracts predetermined instruction information from the reaction data, When the extraction means extracts the instruction information, the control means, when playing back the image data, highlights the object extracted from the content data corresponding to the latest object extracted from the reaction data. The information processing apparatus according to feature 5.

9. A third acquisition means for acquiring the user's facial image data, If the content data includes image data, the derivation means derives position information of the point of focus on the display screen of the display means that displays the image data that the user is looking at, based on the face image data, When the amount of movement of the gaze point per unit time is greater than or equal to a threshold, the control means highlights the object extracted from the content data corresponding to the latest object extracted from the reaction data when playing back the image data. The information processing apparatus according to feature 5.

10. The control means may change the brightness of an object extracted from the content data, change the color of an object extracted from the content data, display a reading ruler on an object extracted from the content data, display a marker on an object extracted from the content data, or delete an object extracted from the content data that corresponds to an object extracted from the reaction data, thereby highlighting the portion of the image data other than the deleted object. The information processing apparatus according to feature 9.

11. The control means changes the playback speed of the content data based on the degree of understanding. The information processing apparatus according to feature 2.

12. The extraction means extracts predetermined instruction information from the reaction data, The information processing apparatus according to any one of claims 1 to 10, characterized in that the control means changes the playback speed of the content data based on the instruction information.

13. A first acquisition step of acquiring content data which includes at least one of image data and audio data, A control step that controls the playback of the content data to be presented by a presentation device including at least one of a display means and an audio output means, A second acquisition step of acquiring reaction data which includes at least one of image data and audio data recording the user's reaction to the content data presented by the presentation device, The process includes an extraction step of extracting object information relating to at least one object of characters, symbols, and figures from the content data and the reaction data, respectively. The control step controls the playback of the content data based on the object information of the content data and the object information of the reaction data obtained in the extraction step. An information processing method characterized by the following:

14. A program for causing a computer to function as an information processing device according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • On-line educating method

    JP2001092340A

  • Moving image playback system and its control method

    JP2009163306A

  • Video processor and video processing method

    JP2013030140A

  • Information processing system

    JP2013105229A

  • Archive system, first terminal, and program

    JP2013114334A