System, apparatus, and method for visually tracking the presentation content of a presenter
The system addresses the challenge of accurately tracking and providing multimedia information in complex environments by using a network of electronic devices and cameras to transmit real-time audio and visual information, enhancing accessibility for impaired individuals.
Patent Information
- Application Number
- JP2022568507
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-17
- Filing Date
- 2021-05-12
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2041-05-12
AI Technical Summary
Current devices struggle to accurately track and provide simultaneous multimedia information, such as the presenter's voice, body language, and written content on a blackboard, especially in complex environments with multiple participants.
A system comprising a first electronic device for the presenter, a second electronic device for the user, a microphone, a fixed camera, a movable camera, a single-board computer, and a router, which provides a short-range communication network to transmit real-time audio and visual information, including presenter tracking data, to ensure universal access to presentation content.
The system enables accurate and universal tracking of presentation content, improving access for both visually and hearing impaired individuals by providing high-quality, real-time multimedia information with minimal delay, even in complex environments.
Smart Images

Figure 0007682510000001 
Figure 0007682510000002 
Figure 0007682510000003
Abstract
Description
Technical Field
[0001] The present invention relates to a system and method for improving the capture and provision of multimedia information, and more particularly, to a system, apparatus, and method for visually tracking the presentation content by a presenter. According to the present invention, it is possible to improve the tracking of presentation content, particularly for hard-of-hearing viewers.
Background Art
[0002] As used herein, the term "presentation" means not only a presentation meeting but also a presentation, a lecture, a meeting, a parliament, an event, an explanation session, an observation in a court, etc. The term "presenter" means not only a presenter in a "presentation" but also a presenter, a teacher, a judge, a lawyer, etc. according to the place. Today, most educational content is provided through conversations, presentations, meetings, etc. These presentations are increasing in number and are attracting more users, enabling viewers to obtain information they desire regarding a specific field. Along with conventional classrooms and their activities, viewers have increasingly higher desires for educational settings, but there are also problems to be overcome. This problem relates to the function of accurately tracking / following classroom activities as universally as possible. That is, it is the ability to adapt to situations / environments where as many people as possible are placed. Solving these problems and overcoming the current difficulties are an obligation and a necessity in today's society, and it is to enable everyone to equally access information.
[0003] From the perspective of viewers participating in a presentation or a meeting, viewers sometimes meet the presenter directly and listen to the speech, observe important things through the body language of the presenter, or may want to know how the presenter is transmitting information. Viewers may also have to look at the blackboard used by the presenter to assist in the presentation or comments. Viewers may also want to see the images, videos, and other information projected from the presentation content through the presenter's computer or the projector in the venue. The combination of these information is necessary for anyone to access in order to correctly follow / understand the presentation / lecture content.
[0004] Current devices cannot reproduce by combining all the information required to fully follow a presentation or a meeting. The main problem with these devices is that it is difficult to transmit multiple pieces of information simultaneously. An example of information includes an image of a blackboard, an image of the presentation content, an image of the presenter moving around, and the voice of the presenter in real time (with a maximum delay of 0.5 seconds). Furthermore, they are facing multiple difficulties. An example of the difficulty is the tracking / capture of the presenter's movement and the capture of the presenter's movement in a situation where there are multiple participants. In this case, this device cannot handle situations where the presenter is facing the blackboard and turning their back to the participants, situations where the presenter is writing something on the blackboard, and situations where the presenter moves instantaneously (a more common situation).
[0005] Another problem is that since current devices cannot perform accurate tracking, they require a very wide shot that can capture the entire blackboard of the scene (sometimes reaching several meters). In this situation, the limitation of current devices is that it is impossible to zoom in to view any part (e.g., the characters written on the blackboard). This is because the image quality required to view the details of the image cannot be maintained.
[0006] In many cases, these devices require a signal generation and processing device to generate and process signals. This signal generation and processing device can provide all the required information in a state suitable for obtaining information about the classroom / presentation venue. However, sometimes these devices must be installed at the venue by professional workers. This is because these devices are not portable.
[0007] These devices are not built-in systems. The reason is that some components of the device are large devices used only by disabled people, and another component is controlled only by one person, and that person is also disabled.
[0008] In recent years, various studies at Spanish universities have been researching and analyzing the possibility of accessing virtual technical environments. What has been found as a conclusion of this research and analysis is that although legal, access to information provided in both the real environment of groups at Spanish universities (where professors give information in the traditional way on the blackboard) and the virtual technical environment (providing digital platforms and materials in electronic format) is quite insufficient, and the technical aids for visually impaired people to access such environments are extremely limited.
[0009] The need for adaptation is the same in other educational environments (e.g., primary schools, middle schools, vocational training schools), and the situation is extremely similar even when analyzed. In traditional classrooms, the means to enable people with visual and hearing impairments to easily access knowledge and information under the same conditions as healthy people are extremely limited.
[0010] Various devices are currently used in educational environments to provide access to and follow the content of all kinds of educational activities, but all of them have some problems.
[0011] One of the currently commercially available systems is the Ablecenter AC-03. The system for visually impaired people has a high-quality 360° camera placed on the ceiling of the room. This camera can be operated by the operator or purchaser and can be manually focused on any place in the room (e.g., the blackboard, the image on the projector). This system can operate on the platforms of TM Android TM , Windows TM , IOS, has a built-in OCR text authentication system and has a function of being able to zoom in on any point in the room. However, this system has the following drawbacks. First, this system is dedicated to visually impaired people, does not emit the voice signal of the speaker, and does not have a voice recognition function, so people with hearing impairments cannot use it. Second, this system can set (focus on) a target at any point in the room, but this is only possible manually, and it cannot track a presenter who moves around freely or recognize the presenter's body language. Third, this system needs to be installed by a professional operator on the ceiling of the room where the camera of this system is intended to be used, and it is not portable. Fourth, the camera of this system can only be controlled by one user, and multiple people in the room cannot select the image on which to focus.
[0012] The current system is Magnilink Sudent Addition. This operation is similar to that of Ablecenter, but the only difference is that the camera is not placed on the ceiling and cannot be controlled by means of a PC. However, the camera can be placed on the desk for the visually impaired and is controlled by the user who directly handles it. The camera can be placed and focused manually. Therefore, this system has the same problems as the above-mentioned one.
[0013] Similarly, there are also cameras with a tracking function (e.g., Lumens VC-TR1, AVer PTC500S). This can track the presenter in a controlled environment, but it is difficult to operate in a complex environment where many people go back and forth with each other. These devices can only collect partial accurate information. Therefore, they cannot cover the above-mentioned requirements. At the same time, it is difficult to execute the tracking function in a complex environment where many people go back and forth with each other.
[0014] [Patent Document 1] US8831505B1 [Patent Document 2] WO2012 / 088443A1 [Patent Document 3] Patent Documents 1 - 3 of US2004 / 002049A1 disclose a conventional system. This conventional system records and broadcasts the content of an educational speech / presentation using a camera and a microphone. This camera and microphone visually capture the information to be transmitted. The signal is sent to a control / generation server via the LAN within the room. Patent Document 1 discloses that tracking of the presenter is performed using a portable device incorporating a wireless microphone and an infrared transmitter. Other Patent Documents 2 and 3 do not disclose a tracking system. These devices have several problems. For example, the control / generation server needs to record, edit, and combine the audio signal and the image signal together with the signal coming from the presenter's PC. It is not possible to achieve a sufficiently small delay time when providing the captured information. Or it is based only on infrared, does not have a presenter tracking system. In a situation where there are multiple people, it cannot be tracked with high reliability all the time with polarized light. The conventional system does not incorporate an access detection tool that adapts and converts access detection information to a general-purpose access detection system.
[0015] Therefore, there is a need to provide a new system, method, and apparatus that enable accurate tracking of the presenter in situ (at the presentation site or classroom) and transfer all information. Examples of information include the information spoken by the presenter, the information shown by actions, the information written on the blackboard, the projector screen information, and the information on the presenter's high-speed PC. These information are for improving the access provided in educational activities.
Summary of the Invention
Problems to be Solved by the Invention
[0016] The object of the present invention is to provide a system, apparatus, and method for tracking the content of an educational activity completely, accurately, accessibly from anywhere, and universally. An example of an educational activity refers to an information presentation activity, a presentation, a class, a meeting, a large meeting, an event, a seminar, etc. Hereinafter, it is collectively referred to as a "presentation" in this specification. The said content is also referred to as "presentation content".
Means for Solving the Problems
[0017] To solve the problems of the prior art, in its first aspect, the present invention provides a system for improving the visual and auditory tracking of the presenter's presentation content. This system includes a first electronic device of the presenter (equipped with first software configured to obtain information within the first electronic device), a second electronic device of the user (equipped with second software), a microphone for acquiring the audio information of the presentation content, a module, and a power source. The module includes a fixed camera for acquiring the information of the presentation content represented on the information display board, a movable camera for continuously acquiring the position information of the presenter, a single-board computer (hereinafter simply referred to as "computer"), and a router for providing a short-range communication network (hereinafter also referred to as "LAN") between the first electronic device and the second electronic device. An example of the information display board is a blackboard, a flip chart, or other means by which the presenter represents the presentation content. The fixed camera, the movable camera, and the single-board computer (hereinafter simply referred to as "computer") are operably connected to the router.
[0018] The system of the present invention has a presenter tracking means (acquiring presenter tracking information based on the position information of the presenter obtained by the movable camera) that can always know the position of the presenter. The second software represents, via the second electronic device, the information within the first electronic device, the audio information of the presentation content, the information shown via the information display board, and the presenter tracking information.
[0019] In one embodiment, the presenter tracking means includes an artificial intelligence algorithm. This artificial intelligence algorithm includes, as an example, a reinforced learning algorithm, a supervised learning algorithm, and an unsupervised learning algorithm. In the system of the present invention, the artificial intelligence algorithm is executed by a computer or a first electronic device. In other embodiments, the system of the present invention includes a remote computer, which is within a cloud computing system (hereinafter simply referred to as "cloud") and is operably connected to the first electronic device, the second electronic device, and the computer. In this case, the artificial intelligence algorithm can be executed by the remote computer device.
[0020] In one embodiment, the system of the present invention further has a recognition mechanism. The recognition mechanism is a band carried by the presenter. An example of the band is a band with at least one of coloring, a logo, and a QR code.
[0021] In one embodiment, the presenter tracking means includes an infrared detection terminal and an infrared detection camera. The infrared detection terminal is carried by the presenter, and the infrared detection camera is connected to the computer.
[0022] In one embodiment, the presenter tracking means has an optical flow tracking algorithm. This algorithm is executed by a computer or a first electronic device
[0023] In one embodiment, in the system of the present invention, the module further has an additional fixed camera. The fixed camera is connected to a router and a power supply.
[0024] The presenter carries the microphone. The microphone captures the presenter's voice and converts it into an electrical signal. This electrical signal is received by a computer or a first electronic device. This signal can be in any signal format. An example thereof is MP3, Windows Media Audio, RIFF, FLV.
[0025] In one embodiment, the system of the present invention further includes a voice receiver. The voice receiver is disposed within the computer or the first electronic device and captures specific other sounds in the room where the presentation is being made.
[0026] In one embodiment, the module further includes an image capture device. The image capture device receives an external image signal directed to the module, the first electronic device, and the second electronic device.
[0027] In one embodiment, the system of the present invention includes a voice recognition device. The voice recognition device converts the voice information (captured by the microphone) during the presentation into a character display. The voice recognition device is disposed within or executed by the computer or the first electronic device. In this case, the computer or the first electronic device transmits the character display to the second electronic device via a LAN.
[0028] In its second aspect, the present invention proposes an apparatus for visually tracking the presentation content of a presenter. This apparatus includes a fixed camera that acquires information on the presentation content presented on an information display board, a movable camera that continuously acquires the position information of the presenter, a presenter tracking means (acquires presenter tracking information based on the position information of the presenter obtained by the movable camera), a router that provides a LAN at the presentation site, a computer, a voice recognition device that converts voice information at the presentation site into a character display, and a power source. The computer, the fixed camera, and the movable camera are operably connected to the router. The computer receives the presenter's information coming from the first electronic device, the voice information of the presentation including the character display, the information appearing on the information display board, and the presenter tracking information, and transmits the received information to the user's second electronic device.
[0029] In one embodiment, the presenter tracking means includes an artificial intelligence algorithm executed by a computer.
[0030] In a third aspect, the present invention provides a method for visually and auditorily tracking the presentation content of a presenter. The method of the present invention includes the following steps (A)-(I). (A) Preparing a module, The module includes a fixed camera, a movable camera, a computer, a router, and a power supply. The fixed camera, the movable camera, and the computer are connected to the router. (B) The router provides a LAN between the first electronic device of the presenter and the second electronic device of the user. (C) Obtaining the information in the first electronic device using first software executed in the first electronic device. (D) Obtaining the audio information of the presentation content with a microphone. (E) Converting the audio information of the presentation content into character display by a speech recognition device. (F) Obtaining the information of the presentation content represented on the information display board using the fixed camera. (G) Continuously obtaining the information of the presenter's position using the movable camera. (H) The presenter tracking means obtains the presenter tracking information based on the presenter's position information obtained in step (G). (I) The second software in the second electronic device represents the information in the first electronic device, the audio information of the presentation content, the information represented on the information display board, and the presenter tracking information.
[0031] In one embodiment, the presenter tracking means includes an artificial intelligence algorithm.
[0032] In one embodiment, the artificial intelligence algorithm is executed on any one of a computer, a first electronic device, and a remote computer in the cloud.
[0033] In one embodiment, the method of the present invention includes the step of obtaining presenter tracking information. The tracking of the presenter is performed based on one of colored markings, logos, characters, and QR codes. All of these are carried by the presenter.
[0034] In one embodiment, the second software transmits and receives all information that appears on either the computer or the first electronic device via the LAN. The LAN is performed using the UDP / multicast communication protocol.
[0035] In one embodiment, the information in the first electronic device, the audio information of the presentation content, the information of the information display board, the presenter tracking information, etc. are sent via the Internet to the electronic devices of users who do not participate in the presentation. These information are transmitted and received via the Real-Time Messaging Protocol (RTMP).
[0036] Therefore, the visual and auditory tracking (understanding) of the presentation content is greatly improved by the present invention. This solution enables the transmission of various signals via the LAN. Simultaneously with the real-time tracking of the presenter moving around in the venue, the following signals are transmitted: the image of the presenter, the image on the board (when the presenter or someone else writes something on the board), the signal of the presentation content transmitted by the first electronic device, the information projected on the board (screen) by the projector, the voice signal of the presenter, etc. Examples of the first electronic device and the second electronic device include a PC, laptop, tablet, electronic tablet, smartphone, mobile phone, etc.
[0037] The second electronic device is within the communication range of the LAN, can receive the above signals with almost no delay, and can transmit them to a predetermined user (listener, student). Therefore, with this solution, the user can follow the presentation content by streaming via their electronic device anywhere in the venue or even outside the venue.
[0038] The present invention is designed so that both visually impaired and hearing impaired people (hereinafter collectively referred to as "visually and hearing impaired people") can access the visual and auditory information of the presentation content that is actually in progress. As a result, according to the present invention, it becomes possible to easily access the presentation content in a large conference hall with many participants, the presentation content in a poorly equipped venue where it is not clearly visible, and the presentation content in a room that is not suitable for all types of audiences.
[0039] All viewers can access the signals transmitted by the system of the present invention by means of a second software (e.g., a computer application). The number of users (students) that the system of the present invention can serve simultaneously can be changed in the basic configuration.
[0040] The device (module of the present invention) of the present invention is light (about 2 kg), small, and portable. The device of the present invention provides great advantages for use in various rooms / places where presenters / professors / speakers go. This device can be quickly and easily started up in various rooms, venues, and conditions, and can be made usable by simply connecting the module to the main socket or outlet. In particular, when visually and hearing impaired people participate, the device of the present invention can be easily applied in a difficult-to-see room, in the case of online recording and transmission.
[0041] In one embodiment, in addition to being able to secure all the information running on the first electronic device, the first software (i.e., the master software) of the first electronic device can further transmit the remaining signals. The second software of the second electronic device enables the reception of such signals. The tracking software is installed on a computer or in the cloud, enables the transmission of all signals, and further enables the tracking of the presenter.
[0042] In one embodiment, the first software is available on various platforms (PC, IOS TM , Android TM) It is operable. The first software enables transmission over a LAN. This is achieved through the streaming of three types of image signals and an audio signal, allowing for the recording of a complete session and / or the separate recording of images and audio. The three types of image signals are the image of the presenter, the image of the board, and the image of the screen of the presenter's first electronic device. Streaming is limited to cases where a wireless internet connection is available.
[0043] The second software is developed to be installed on PCs, IOS TM , Android TM The second software enables the electronic device to receive three types of image signals and an audio signal. In the second software, these signals can be changed according to requirements and can be viewed on a pair (or multiple) of displays or on a single display, with a screen change delay of 0.5 seconds or less. Zooming is possible for any image, and the contrast (black on white or vice versa) can also be changed. Additionally possible are stopping or resuming any signal, acquiring a screen shot (still image), recording it to an archive, and playing back the recorded images and audio signals. These are performed in a complete session with the software of the presenter's first electronic device, making it seem as if one were at the venue.
[0044] Thus, the system of the present invention is simple and practical, allowing for inexpensive and efficient operation without the modules overheating. As a result, a large number of users can connect to the system of the present invention simultaneously. This is done smoothly without interruption and can cover the entire presentation content with high image quality.
Brief Description of the Drawings
[0045]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Mode for Carrying Out the Invention
[0046] The present invention proposes a system, apparatus, and method that enable a third party to easily track the content of a presenter's presentation (by visual, auditory, text display, etc.). In particular, a system, apparatus, and method are proposed that enable a person with visual or hearing impairments (hereinafter referred to as a "visually / audiologically impaired person") to access the knowledge and information of the presentation content under the same conditions as a healthy person.
[0047] FIG. 1 shows an embodiment of the system of the present invention. With this system, visual / auditory information of the presentation content is acquired, and the obtained information is transmitted in real time (delay time of 0.5 seconds or less) to a second electronic device 2 via a short-range wireless communication network (hereinafter referred to as a "LAN"). Therefore, according to this embodiment, this system includes a first electronic device 1 of the presenter, a second electronic device 2 of the user, i.e., an attendee of the presentation (hereinafter collectively referred to as the "user"), a microphone 3, and a module 5. The microphone 3 is a device that converts sound into an electrical signal (regardless of analog or digital).
[0048] In this embodiment, module 5 includes a fixed camera 8, a movable camera 9, a single-board computer (hereinafter simply referred to as "computer") 10, a router 6, and a power supply. An example of the power supply is a 12V power supply. Examples of the fixed camera 8 and the movable camera 9 include digital / electronic cameras, video cameras, 2D cameras, and 3D cameras, which capture images in analog or digital format. The field of view of the fixed camera 8 focuses on a display board (hereinafter simply referred to as "board") 28 and reads the information displayed thereon. The movable camera 9 continuously acquires the position information of the presenter. The rotation of the movable camera 9 is performed by means 21 for moving the movable camera 9 (Figs. 3A - 3C). The field of view of the movable camera 9 is directed towards the direction of the presenter's position. As a result, the movable camera 9 is always focused on the presenter. The router 6 provides a LAN between the presenter's first electronic device 1 and the user's second electronic device 2.
[0049] As shown in Fig. 1, module 5 is arranged in the room where the presentation is conducted. Thereby, the camera system of module 5 focuses on locations away from the board 28, the projection area, the presenter, and the first electronic device 1. An example of the first electronic device 1 is a PC, laptop, tablet, mobile phone, etc., and it can be any means by which the presenter discloses information. The connection is made via a LAN. Since module 5 is fixedly arranged, it starts operating just by being connected to the LAN.
[0050] Similarly, this system has a presenter tracking means. This presenter tracking means obtains presenter tracking information based on the position information of the presenter obtained by the movable camera 9. In one embodiment, the presenter tracking means has an artificial intelligence algorithm. In one example, this artificial intelligence algorithm detects the texture of the segmented image of the presenter. In other embodiments, the presenter tracking means is a combined system of an artificial intelligence algorithm and an authentication mechanism / element. An example of the authentication mechanism / element is a band with color markings, logos, or QR codes. The use of the presenter tracking means can prevent tracking failures. Examples of the causes of tracking failures include when the presenter turns their back to the camera and in dim situations.
[0051] In other embodiments, the presenter tracking means includes an infrared camera 7 connected to the computer 10 and an infrared detection terminal carried by the presenter.
[0052] Furthermore, in FIG. 1, both the presenter's first electronic device 1 and the user's second electronic device 2 each have software (or communication management means) to obtain predetermined information. The predetermined information includes information running on the first electronic device 1, information running on the second electronic device 2, audio information obtained by the microphone 3, information displayed via the board 28, and presenter tracking information. The software executed within the first electronic device 1 is referred to as first software, which receives and displays the predetermined information. The software executed within the second electronic device 2 is referred to as second software, which obtains information running on the first electronic device 1 in addition to the above-mentioned predetermined information. Similarly, the computer 10 has software. This software controls the reception of signals / information obtained from the camera, the microphone 3, and the presenter tracking means, and controls the transmission of signals / information to the second electronic device 2. An example of this camera is the fixed camera 8, the movable camera 9, and an external camera (if any).
[0053] Therefore, the signals recorded and transmitted by the system of the present invention are (1) The first shot and real-time image of the presenter while continuously tracking the presenter. Thereby, hear the presenter's voice and feel the gestures. (2) The image of the board 28 used by the presenter. Thereby, it is possible to see what the presenter has written / drawn on the board 28. (3) The presentation content, video, and images of the problems raised in the presentation shown on the first electronic device 1. (4) The real-time audio signal of the professor's speech. In one embodiment, the system of the present invention further (5) Also acquires images by other cameras by the image capture element executed within the module 5.
[0054] The second software executed within the user's second electronic device 2 ensures communication with the corresponding communication device, particularly through the UDP / multicast communication protocol. This protocol can improve connectivity among a vast number of users.
[0055] In addition to the above LAN, signals can be transmitted via the online (e.g., the Internet, for example, the Real Time Messaging Protocol) to other users who do not participate in the presentation.
[0056] Particularly, the transmission of image signals via the LAN utilizes the H.264 image compression format. This compression format involves a minor loss of information without filters, but it can increase the transmission speed while maintaining the quality of the final image by reducing the compression size. Other image compression formats can also be used.
[0057] Therefore, module 5 operates on the LAN without the need for the Internet. When module 5 has to transmit signals via the Internet to those who do not participate in the presentation, module 5 has a communication device or communication module (4G, 5G card). When the presentation venue does not have Internet services or the radio waves there are weak, module 5 can provide Internet services to this system.
[0058] In one embodiment, module 5 further has an additional fixed camera 8.1 (Figure 2). The additional fixed camera 8.1 is connected to the router 6 and the power supply 13. The additional fixed camera 8.1 assists the fixed camera 8. The additional fixed camera 8.1 and the fixed camera 8 are arranged so that they cover all specific locations in the presentation venue. With the additional fixed camera 8.1, the entire board 28 can be photographed with high quality and high clarity. The additional fixed camera 8.1 can focus on the entire presentation venue including the board 28 or the projection area. It can also focus on other specific locations in the presentation venue. An example of other specific locations is a board with additional information, a signal for the narrator, etc. With the signal for the narrator, even those with hearing impairments can understand the explanation.
[0059] The presence of the additional fixed camera 8.1 and the fixed camera 8 has advantages compared to the current system where there is only one camera and it cannot cover the space of the presentation venue, for example, the entire blackboard. In fact, a wide-angle camera cannot cover as wide a range as initially thought. The reason is that when the field of view is significantly widened as an image collection mode, the resolution of a small area deteriorates significantly. Thus, when zooming in to enlarge the words or text on the board 28, they become less clearly visible. However, this is different when using a large, high-resolution, professional-grade wide-angle camera (not suitable for solving the present invention). When using the additional fixed camera 8.1, especially when using the same type as the fixed camera 8, the wide field of view or display area of the board 28 can be enlarged, but even in this case, the image is not distorted and does not lose clarity.
[0060] Similarly, in one embodiment, the module 5 has a voice recognition device. This voice recognition device is connected to the first electronic device 1 and the computer 10. With this voice recognition device, the module 5 can transcribe the speaker's speech in real time in the form of subtitles in a predetermined language. In this case, a device capable of transcribing voice into text in real time is required. A device for recognizing the text of the OCR of the image of the presentation content is activated, and the voice signal of the speech is input to the voice assistance means for the hearing-impaired. In other embodiments, the voice recognition device may be mounted on the first electronic device 1.
[0061] In one embodiment, when power is applied, module 5 automatically operates. That is, it receives predetermined signals / information without passing all information through the control device and the production center. Examples of the predetermined signals / information include the image of the presenter, optionally the images around the presenter, the fixed image of board 28, the signal from the first electronic device 1, the audio signal, and optionally the external image signal obtained by the image capture device 14. The signal / information from the first electronic device 1 is received and transmitted to the second electronic device 2 only at any time or when the presenter activates its function (sharing / double-transmitting its computer screen with board 28). For example, the software asks the presenter via the user interface of the presenter's own first electronic device whether the presenter wants to share his or her screen. If the answer is YES, the signal of the content of the screen of the first electronic device 1 is sent to the computer 10. The same applies to the remaining signals. That is, the fixed signal of board 28, the image of the presenter, and the audio signal are received by module 5 and sent to the second electronic device 2 constantly or when the corresponding function is activated. Therefore, when the signal is blocked (for image protection, cost reduction, or power consumption reduction of the device), module 5 operates by transmitting only one signal. For example, with an appropriate configuration of the device, the device can transmit only the signal of the presenter's computer screen and the audio signal. For this purpose, the system administrator or the presenter himself / herself can contact module 5 via an appropriate interface and block any signal, that is, not transmit it. Therefore, the system administrator (or the presenter himself / herself) can interact with module 5 via an appropriate interface and block any signal, that is, not transmit it.
[0062] In another embodiment, the first software running on the first electronic device 1 is edited to receive the above-mentioned signals / information. In this case, the first software is configured to send all these signals / information to the router 6 and then to the second electronic device 2 from there. The communication information is transmitted to the second electronic device 2 at any time or when the presenter activates its function (sharing the screen of the first electronic device 1 with board 28 / displaying on both).
[0063] In yet another embodiment, the above-mentioned signal / information can be received by a remote computer. This remote computer is located in the cloud (not shown) and is operably connected to the first electronic device 1, the second electronic device 2, and the computer 10.
[0064] FIG. 2 shows an example of the connection state among various devices included in module 5. In this example, the movable camera 9, the fixed camera 8, the additional fixed camera 8.1, and the router 6 are powered by the power supply 13. At the same time, it has a second power supply 13.1 (e.g., a 5V power supply), which serves as the power supply for the computer 10. In this example, module 5 has an infrared camera 7. The infrared camera 7 is directly supplied with current from the computer 10 via a USB port. An example of the computer 10 is Jetson Nano or Raspberry pi. The power line 11 is represented by a dotted line in FIG. 2. The data connection between the fixed camera 8, the additional fixed camera 8.1, the movable camera 9, and the router 6 is made by two Ethernet cables 12 (shown as solid lines). There are two data connections between the computer 10 and the router 6. The first data connection provides an Internet service to the router 6. This is the first function of the computer 10. The Ethernet port of the router 6 is an Internet input port. The second data connection enables communication between the movable camera 9 and the computer 10 and enables tracking. This is the second function of the computer 10. This cable comes out of the USB outlet of the computer 10 and enters the router 6 with an Ethernet outlet. As a result, a coupling that enables switching is required. This is because the computer 10 does not have two Ethernet input jacks.
[0065] The movable camera 9 is preferably built into the module 5. However, in one embodiment, the movable camera 9 can be separated from the module 5 and communicate between them by wire or wirelessly. The system of the present invention may have a plurality of movable cameras 9.
[0066] When the infrared camera 7 is incorporated, the infrared camera 7 can also communicate with the computer 10 via a USB cable.
[0067] In one embodiment, the presenter tracking means has an optical flow tracking algorithm. This algorithm is executed on the computer 10 or the first electronic device 1. In one embodiment, the optical flow tracking algorithm operates when a failure or interruption occurs in at least one of the infrared cameras 7. Infrared detection is performed by the infrared camera 7 and the infrared detection terminal disposed in the module 5. This infrared detection terminal is carried by the presenter, emits infrared rays, and the infrared camera 7 detects it. As a result, the position of the presenter is detected, and an image of the presenter is captured by the movable camera 9. In this embodiment, the microphone 3 is also carried by the presenter, built into the portable infrared detection terminal, and captures and transmits the voice signal of the presenter.
[0068] In one embodiment, infrared tracking is mainly used. For some reason, for example, when the presenter turns his back to the camera when writing on the board 28, when interfering with other infrared rays in the venue, when losing the infrared signal, or when a collision with other infrared signals occurs, the tracking by optical flow automatically operates. The optical flow tracking means / algorithm starts automatically or is controlled by the computer 10 or manually by the presenter. In this sense, the presenter can also select the first software to perform tracking by optical flow detection instead of infrared detection. This first software communicates with the computer 10 and sends execution instructions to the computer 10. In the case of tracking by the optical flow algorithm, the protocol identifies the movement in the image by comparing a frame at a certain point in the used image sequence with the subsequent one. The speed of the movement of the camera is fixed based on the change in the position of the presenter detected in the consecutive images of the video.
[0069] In the present invention, having an infrared camera 7 and an optical flow tracking algorithm is an optional matter. For example, in an embodiment where the presenter tracking means is executed by the above algorithm, an artificial intelligence algorithm, or a combined system of both, both the infrared camera 7 and the optical flow tracking algorithm are not necessarily required.
[0070] In one embodiment, the algorithm is based on the OpenCV computer vision library. It uses a pre-trained multi-object detection model suitable for detecting people. An example of this model is the Google TM model. Further additional processing can be added to train based on the first person passing through the target position. In one embodiment, the detection can be performed using a comparison of the HSV (Hue, Saturation, Value) histogram and the shape format of the detected person, or by a segmentation method. When the person to be tracked (e.g., the presenter) is identified, the position change of the person in the consecutive images sent by the movable camera 9 that captures the presenter's image to is determined and sent to the port and IP address of the movable camera 9 by means of a TCP connection via a command (indicating where the presenter is). Incidentally, the movable camera 9 has a port for 16-digit control commands of the VISCA protocol. The speed of the movement of the camera is fixed based on the change in the position of the presenter detected in the consecutive images of the video. Other computer vision mechanisms for tracking people can also be used.
[0071] Regarding tracking by QR code recognition means, the QR code (tag, sticker, attached to a mobile device, etc.) carried by the presenter is used as a marking, and the position and distance of the presenter are identified by constantly detecting that marking. For this reason, in one embodiment, the OpenCV computer vision library, particularly a QR type plate detector, is reused. The movable camera 9 is calibrated with a marking of a known size on the surface to obtain its intrinsic parameters as a matrix of focal length and distortion. With this information and the marking of a known size, the marking is detected and identified by detecting characteristic points. Therefore, if the presenter has a predetermined marking (QR code), the position and distance of the person (presenter) having that marking are constantly identified.
[0072] Regarding tracking by color marking, the color marking (tag, sticker, attached to a mobile device, etc.) carried by the presenter is used as a marking, and the position and distance of the presenter are identified by constantly detecting that marking. This protocol filters the color of the image captured by the movable camera 9 to identify the position of the person to be tracked.
[0073] The above description relates to one embodiment, and other mechanisms for tracking a person by means of marking can also be used.
[0074] By combining these tracking means (means such as artificial intelligence, QR code authentication, color marking, etc.), malfunctions or interruptions that may occur when the presenter turns their back to the camera during tracking or when the venue is dim are eliminated. Accurate tracking of the presenter is important because by constantly receiving an accurate image of the presenter, the receiver can sense the presenter's body language (gestures). This is extremely important in information transmission. Furthermore, in this tracking mechanism / means, it is possible to reproduce with higher resolution compared to the movable camera 9 (which focuses on the entire board 28). With this configuration, not only the normal shots on the board 28 (by the fixed cameras 8 and the additional fixed camera 8.1), but also the reduced shots of what the presenter is writing at a given point in time by the movable camera 9 (by the movable camera 9) can be obtained with higher quality. As a result, the zoom function is improved.
[0075] Figures 3A - 3C show an embodiment of the module 5 equipped with the fixed camera 8 and the movable camera 9. In Figure 3C, the ON / OFF button 15 of the module 5, the connector 16 that supplies power to the module 5, the RJ45 connector 17, and the HDMI connector 18 are shown. In one embodiment, the module 5 can be remotely operated.
[0076] The above description relates to an embodiment of the present invention. Those skilled in this technical field can conceive various modifications of the present invention, all of which are included in the technical scope of the present invention. The numbers in parentheses described after the components of the claims correspond to the component numbers in the drawings and are attached for the easy understanding of the invention, and are not intended to limit the interpretation of the invention (Article 24, Paragraph 4 and Form 29, Paragraph 2, "Remarks" 14 (b) of the Implementing Regulations of the Patent Law). Also, even if the same number is used, the component names in the specification and the claims are not necessarily the same. This is due to the reasons described above. "At least one or more" and "and / or" are not limited to one of them. For example, "at least one of A, B, and C" may include not only "A", "B", and "C" alone but also a plurality of them such as "A, B or B, C or further A, B, C". In this specification, "including A" and "having A" may include things other than A. Unless otherwise specified, the number of devices or means may be singular or plural. The present invention can be applied to any "place of publication" by any "person who makes a publication". An example of the "place of publication" is a conference venue, a school classroom, a meeting, an event, a seminar, a court, etc.
Explanation of Signs
[0077] 1: First electronic device 2: Second electronic device 3: Microphone 5: Module 6: Router 7: Infrared camera 8: Fixed camera 8.1: Fixed camera 9: Movable camera 10: Computer 11: Power line 12: Ethernet cable 13: Power supply 13.1: Second power supply 15: Switch 16: Connector 17: RJ45 connector 18: HDMI connector 28: Information display board
Claims
1. In a system for visually and auditorily tracking the presentation content of a presenter, (A) a first electronic device (1) of the presenter, (B) a second electronic device (2) of the user, (C) a microphone (3) for acquiring audio information of the presentation content, (D) a portable module (5), (E) presenter tracking means, (F) having a voice recognition device, The first electronic device (1) is equipped with first software configured to obtain information within the first electronic device (1), The second electronic device (2) is equipped with second software, The module (5) includes: (D1) a fixed camera (8) for acquiring information on the presentation venue represented on an information display board (28), (D2) a movable camera (9) for continuously acquiring the position information of the presenter, (D3) a computer (10), (D4) a router (6) that provides a short-range communication network (hereinafter referred to as "LAN") between the first electronic device (1) and the second electronic device (2), (D5) having a power source, The presenter tracking means (E) obtains tracking information of the presenter based on the position information of the presenter obtained by the movable camera (9), The voice recognition device (F) is stored in the first electronic device (1) or the module (5) and converts the voice information obtained by the microphone (3) into character display, The first electronic device (1) or the computer (10) transmits the character display to the second electronic device (2) via the LAN, The fixed camera (8), the movable camera (9), and the computer (10) are operably connected to the router (6), The second software is configured to display, via the second electronic device (2), the information executed by the first electronic device (1), the audio information of the presentation including the character display, the information displayed via the information display board (28), and the tracking information of the presenter. A system for visually and auditorily tracking the presentation content of a presenter, characterized by the above.
2. The presenter tracking means (E) includes an artificial intelligence algorithm. The system according to Claim 1, characterized by the above.
3. Further having a recognition mechanism carried by the presenter, the recognition mechanism recognizing one of color, logo, and QR code. The system according to Claim 2, characterized by the above.
4. The presenter tracking means (E) includes: an infrared detection terminal carried by the presenter and an infrared detection camera connected to the computer (10), or Having an optical flow tracking algorithm executed by the computer (10) or the first electronic device (1) The system according to claim 1, characterized in that.
5. The artificial intelligence algorithm is executed by the computer (10) or the first electronic device (1) The system according to claim 2, characterized in that.
6. Further comprising a remote computer device, the remote computer device being arranged within a cloud computing system and operably connected to the first electronic device (1), the second electronic device (2) and the computer (10), The artificial intelligence algorithm is executed by the remote computer device The system according to claim 2, characterized in that.
7. The module (5) further has an additional fixed camera (8.1), and the fixed camera (8.1) is connected to the router (6) and the power supply (13) The system according to claim 1, characterized in that.
8. The microphone (3) is carried by the presenter The system according to claim 1, characterized in that.
9. Further having a voice receiver, the voice receiver being connected to the computer (10) or the first electronic device (1) and capturing other sounds in the room where the presenter is located The system according to claim 8, characterized in that.
10. The module (5) further has an image capture device, and the image capture device receives an external image signal directed to the module (5), the first electronic device (1) and the second electronic device (2) The system according to claim 1, characterized in that.
11. In a portable device (5) for visually tracking the presentation content of a presenter, The portable device (5) is,[[]] (A) A fixed camera (8) for acquiring information on the presentation content represented on the information display board (28), (B) A movable camera (9) for continuously acquiring the position information of the presenter, (C) A presenter tracking means, (D) A router (6) for providing a LAN at the presentation site, (E) A computer (10), (F) A voice recognition device for converting voice information at the presentation site into character display, (G) Having a power supply, The presenter tracking means obtains tracking information of the presenter based on the position information of the presenter obtained by the movable camera (9), The computer (10), the fixed camera (8) and the movable camera (9) are operably connected to the router (6), The computer (10) is,[[]] Receive the information of the presenter from the first electronic device (1), the audio information of the presentation including text display, the information appearing on the information display board (28), and the tracking information of the presenter. Transmit the received tracking information of the presenter to the second electronic device (2) of the user. A portable device for visually and auditorily tracking the presentation content of a presenter, characterized by the above.
12. The presenter tracking means includes an artificial intelligence algorithm executed by the computer (10). The portable device according to claim 11, characterized by the above.
13. In a method for visually and auditorily tracking the presentation content of a presenter, (A) Prepare a module (5) which is a portable device. The module (5) has a fixed camera (8), a movable camera (9), a computer (10), a router (6), presenter tracking means, and a power supply. The fixed camera (8), the movable camera (9), and the computer (10) are connected to the router (6). The module (5) has a fixed camera (8), a movable camera (9), a computer (10), a router (6), presenter tracking means, and a power supply. The fixed camera (8), the movable camera (9), and the computer (10) are connected to the router (6). (B) The router (6) provides a LAN between the first electronic device (1) of the presenter and the second electronic device (2) of the user. (C) Obtain the information in the first electronic device (1) using the first software executed in the first electronic device (1). (D) Obtain the audio information of the presentation content with a microphone (3). (E) Convert the audio information of the presentation content into text display by an audio recognition device. (F) Obtain the information of the presentation content represented on the information display board (28) with the fixed camera (8). (G) Continuously obtain the information of the position of the presenter with the movable camera (9). (H) The presenter tracking means obtains the tracking information of the presenter based on the information of the position of the presenter obtained in step (G). (I) The second software in the second electronic device (2) represents the information in the first electronic device (1), the audio information of the presentation including text display, the information represented on the information display board (28), and the tracking information of the presenter. Have A method for visually and auditorily tracking the presentation content of a presenter, characterized by the above.
14. The presenter tracking means is an artificial intelligence algorithm. The artificial intelligence algorithm is executed on any one of the computer (10), the first electronic device (1), and the remote computer of the cloud computing system. The method according to claim 13, characterized by the above.
15. The second software receives all information appearing in either the computer (10) or the first electronic device (1) via a LAN. The information within the first electronic device (1), the voice information of the presentation content, the information on the information display board (28), and the tracking information of the presenter are sent via the Internet to the electronic devices of users who do not participate in the presentation venue. The method according to claim 13, characterized in that.
Citation Information
Patent Citations
Photographing device
JP1997018849A
Image processing unit
JP2000350192A
Presentation-image distribution system
JP2010087613A
System and method for automated capture and compaction of instructional performances
US20130285909A1
Automatic tracking camera control system
US5434617A