Video display system, video display device, and video display method

The video display system accurately visualizes sounds and objects in real-time by integrating object and sound detection, addressing the mismatch in existing technologies and enhancing the immersive experience for spectators, especially in sports venues.

JP2026001822APending Publication Date: 2026-01-08AISIN CORP +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024099347
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing technologies fail to accurately visualize sounds in moving images, particularly in sports venues, as they do not determine the actual sound corresponding to audio information, leading to mismatched onomatopoeia.

Method used

A video display system that includes a photographing unit, object detection unit, sound collection unit, sound detection unit, display object determination unit, and display control unit, which work together to detect objects and sounds in real-time, determining display objects corresponding to the actual situation at the venue based on captured images and sound information.

Benefits of technology

Enables accurate visualization of objects and sounds in real-time, allowing spectators to experience the actual venue situation, including those with hearing impairments, through flexible display formats and immersive experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026001822000001_ABST
    Figure 2026001822000001_ABST
Patent Text Reader

Abstract

To display a display object corresponding to an actual situation of a hall with respect to a photographed image on the basis of the photographed image and sound information acquired in the hall.SOLUTION: The object and advantages of the invention will be realized and attained by means of the elements and combinations particularly pointed out in the claims. According to an aspect of the invention, there is provided a video display system including a photographing unit that photographs a venue, a target object detection unit that detects a predetermined target object based on a photographed image acquired from the photographing unit, a sound collection unit that collects sound in the venue, a sound detection unit that detects sound in the venue based on sound information acquired from the sound collection unit, a display object determination unit that determines a display object corresponding to a situation in the venue based on a detection result of the target object detection unit and a detection result of the sound detection unit, a display unit that displays information, and a display control unit that displays the photographed image and the display object on the display unit.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video display system, a video display device, and a video display method. [Background technology]

[0002] Conventionally, there have been techniques for inserting marks, characters, etc. into photographed images or videos. For example, there is a technique (hereinafter referred to as Prior Art 1) that estimates an object in a photograph by image recognition based on photographed data, and inserts onomatopoeic characters (onomatopoeic words, mimetic words, and words with an imitative voice) corresponding to the estimated object into the displayed image, or inserts non-characters such as marks and lines.

[0003] Prior art 1 also records surrounding sounds when taking a photo and stores them in association with the image file, and can insert strings such as "scene," "noisy," or "bustling" from the sound data linked to the image file to indicate the situation based on the volume.

[0004] There is also a technology (hereinafter referred to as Prior Art 2) that evaluates the surrounding environment from the measurement results of a measuring device, acquires onomatopoeia corresponding to the evaluation content, and notifies the user with images and sounds. In Prior Art 2, the surrounding environment is evaluated by, for example, using the speed at which trees sway recognized from the image analysis results, using the diameter of stones lying on the road, analyzing the size of raindrops, or calculating the degree of quietness from audio data obtained by an audio sensor. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-205296 [Patent Document 2] Patent No. 6917311 Summary of the Invention [Problem to be solved by the invention]

[0006] However, when considering the visualization of sounds in photographed images of a venue (such as a sports venue), for example, the above-mentioned prior arts 1 and 2 are not sufficient. Specifically, prior art 1 is aimed at decorating still images, and does not describe a method that can be used with moving images. Furthermore, because prior art 1 does not determine what sound a given sound is from the audio information, if onomatopoeia is simply estimated from objects in the image, it is conceivable that the onomatopoeia displayed for a sound may not match the actual sound.

[0007] Furthermore, in prior art 2, audio data is only used to determine volume information, so as with the above, no determination is made as to what sound it is, and it is possible that the onomatopoeia displayed for a sound may not match the actual sound.

[0008] The present invention has been made in consideration of the above circumstances, and aims to provide a video display system, a video display device, and a video display method that can display an object corresponding to the actual situation at the venue in relation to a captured image based on the captured image and sound information acquired at the venue. [Means for solving the problem]

[0009] The video display system of this embodiment comprises a photographing unit that photographs the venue, an object detection unit that detects a specified object based on the photographed image acquired from the photographing unit, a sound collection unit that collects sounds from the venue, a sound detection unit that detects sounds from the venue based on sound information acquired from the sound collection unit, a display object determination unit that determines a display object corresponding to the situation of the venue based on the detection results by the object detection unit and the detection results by the sound detection unit, a display unit that displays information, and a display control unit that displays the photographed image and the display object on the display unit.

[0010] The video display device of this embodiment includes an object detection unit that detects a specified object based on a captured image obtained from a capture unit that captures images of the venue, a sound detection unit that detects sounds in the venue based on sound information obtained from a sound collection unit that collects sounds in the venue, a display object determination unit that determines a display object corresponding to the situation of the venue based on the detection results by the object detection unit and the detection results by the sound detection unit, and a display control unit that displays the captured image and the display object on a display unit that displays information.

[0011] The video display method of this embodiment includes an object detection step in which an object detection unit detects a specified object based on a captured image acquired from a capture unit that captures images of the venue; a sound detection step in which a sound detection unit detects sounds of the venue based on sound information acquired from a sound collection unit that collects sounds of the venue; a display object determination step in which a display object determination unit determines a display object corresponding to the situation of the venue based on the detection result by the object detection unit and the detection result by the sound detection unit; and a display control step in which a display control unit displays the captured image and the display object on a display unit that displays information. [Effects of the Invention]

[0012] According to this embodiment, it is possible to display an object corresponding to the actual situation of the venue on the basis of the photographed image and sound information acquired at the venue. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a diagram showing an outline of the overall configuration of a video display system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating a functional configuration of the information processing apparatus according to the embodiment. [Figure 3] FIG. 3 is a diagram showing an example of a display screen in the embodiment. [Figure 4] FIG. 4 is a table relating to a method for displaying onomatopoeia in the embodiment. [Figure 5] FIG. 5 is a flowchart illustrating processing by the information processing apparatus according to the embodiment. [Figure 6] FIG. 6 is a flowchart showing the details of the process in step S5 of FIG. [Figure 7] FIG. 7 is a flowchart showing the details of the process in step S7 of FIG. [Figure 8] FIG. 8 is a flowchart showing the details of the process in step S9 of FIG. DETAILED DESCRIPTION OF THE INVENTION

[0014] The video display system, video display device, and video display method of this embodiment will be described below with reference to the drawings. In this embodiment, a competition venue where a table tennis match is held is used as an example, and the situation, atmosphere, cheering, etc. of the venue during the match are visualized in real time. This allows everyone to enjoy the match, regardless of whether the spectators are at the competition venue or in a remote location, and regardless of whether the spectators are hearing-impaired (those with normal hearing). Such technology is proposed.

[0015] 1 is a diagram showing an outline of the overall configuration of a video display system S according to an embodiment. The video display system S includes a camera 2 (imaging device), a microphone 3 (sound collecting unit), an information processing device 4, and a venue screen 5 (first display device) that are arranged in a venue V (competition venue) where a ping-pong table 1 is installed, and a smartphone 6 (second display device), a PC (Personal Computer) 7 (second display device), and VR (Virtual Reality) goggles 8 (second display device) that are arranged in a remote location R and connected to the information processing device 4 (video display device) via the Internet via a cloud server C.

[0016] Camera 2 is an example of a camera unit that captures images of venue V. Although only one camera 2 is shown in FIG. 1, there may be multiple cameras. In addition to a normal camera, camera 2 may be, for example, a depth camera or a stereo camera that can measure distance. A number and types of cameras 2 are installed in venue V that can collect the video information necessary to recognize the current actual situation of venue V.

[0017] Microphone 3 is an example of a sound collection unit that collects sounds in venue V. Although only one microphone 3 is shown in FIG. 1, there may be multiple microphones. In addition to a normal microphone, the microphone 3 may also be, for example, a vibration sensor that detects vibrations and outputs the amplitude of the vibrations. A number and types of microphones 3 are installed in venue V that can collect the sound information necessary to recognize the current actual situation of venue V.

[0018] The venue screen 5 is a display device that can be viewed by spectators in the venue V.

[0019] The smartphone 6 is used by a user at a remote location R to display images relating to the competition at the venue V.

[0020] PC7 is used by a user at a remote location R to display images relating to the competition at venue V.

[0021] The VR goggles 8 are worn and used by a user in a remote location R, and display images related to the competition at the venue V, for example, in the metaverse space.

[0022] Next, the functional configuration of the information processing device 4 will be described with reference to Fig. 2. Fig. 2 is a diagram showing the functional configuration of the information processing device 4 according to the embodiment. The information processing device 4 is a computer device, and includes a processing unit 41, a storage unit 42, an input unit 43, a display unit 44, and a communication unit 45.

[0023] The storage unit 42 is realized by, for example, a random access memory (RAM), a read only memory (ROM), a solid state drive (SSD), a hard disk drive (HDD), etc., and stores various types of information. The storage unit 42 stores, for example, the operation program of the processing unit 41, various types of data, various types of calculation results, etc.

[0024] The input unit 43 is a means for the user to input information, and is realized by, for example, a keyboard or a mouse.

[0025] The display unit 44 is a means for displaying information, and is realized by, for example, an LCD (Liquid Crystal Display).

[0026] The communication unit 45 is a communication interface for communicating with an external device (such as the cloud server C).

[0027] The processing unit 41 is realized by, for example, a CPU (Central Processing Unit) and executes various information processes. The processing unit 41 includes, for example, as functional components, an acquisition unit 411, an object detection unit 412, a sound detection unit 413, an object action classification unit 414, a sound classification unit 415, a display object determination unit 416, a display position determination unit 417, an association unit 418, a display form determination unit 419, and a control unit 420 (display control unit).

[0028] To facilitate understanding of the following description, an example of a display screen in the embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of a display screen in the embodiment.

[0029] 3, the image captured by camera 2 shows a table tennis table T, a table tennis ball B, player A1 and his racket RA1, player A2 and his racket RA2, and spectator seats. Objects D1 to D4 are superimposed on the captured image to visualize the situation, atmosphere, cheering, and the like during the match in real time. Details will be described later.

[0030] 1 and 2, the acquisition unit 411 acquires various types of information. For example, the acquisition unit 411 acquires a captured image (video information) from the camera 2 and acquires sound information from the microphone 3.

[0031] The object detection unit 412 detects a predetermined object based on the captured image acquired from the camera 2. In this embodiment, a table tennis match using a ball is played by players (hereinafter also referred to as "players") in front of spectators at a competition venue. In this case, the predetermined object includes, for example, at least one of the players, the ball, and the spectators. The object detection unit 412 detects, for example, the position and movement of the object based on the captured image.

[0032] The sound detection unit 413 detects sounds in the venue based on sound information acquired from the microphone 3 .

[0033] The object motion classification unit 414 classifies the motion of the object based on the detection result by the object detection unit 412. For example, if the object is a player, the object motion classification unit 414 classifies the motion of the player using a classification model (e.g., DDnet) for classifying human motion by tracking joint points based on the detection result by the object detection unit 412. The object motion classification unit 414 may also determine the posture of the player based on the detection result by the object detection unit 412.

[0034] Based on the detection results from the sound detection unit 413, the sound classification unit 415 classifies the detected sounds into sounds belonging to the venue V (e.g., sounds of applause and cheers from spectators) or sounds belonging to an object (e.g., sounds of a ball and sounds of players' shoes). For example, the sound classification unit 415 classifies the detected sounds based on a predetermined learning model or a transfer learning model generated by transfer learning using the predetermined learning model. For example, sounds of applause and cheers from spectators can be classified by a predetermined learning model (e.g., public models YAMNet and RandomForest). Also, for example, sounds of a ball and sounds of players' shoes can be classified by a transfer learning model (e.g., a transfer learning model based on YAMNet).

[0035] The display object determination unit 416 determines a display object corresponding to the situation of the venue V based on the detection result by the object detection unit 412 and the detection result by the sound detection unit 413. The display object determination unit 416 may also determine a display object corresponding to the situation of the venue V based on the classification result by the object action classification unit 414 and the classification result by the sound classification unit 415.

[0036] The display position determination unit 417 determines the display position of the object based on the detection result by the object detection unit 412. For example, if the object is a ball, the display position determination unit 417 determines the display position of the object corresponding to the ball by estimating the position of the ball using the difference in color space between the ball and the ping-pong table. For example, in the display screen example of FIG. 3, the display position determination unit 417 determines the display positions of the "span" of the object D2 corresponding to the ball and the "catch" of the object D3 to be near the ball B.

[0037] The association unit 418 and the display form determination unit 419 will be described later.

[0038] The control unit 420 executes various controls. The control unit 420 (display control unit) displays the captured image and display objects on a display unit (such as the venue screen 5, smartphone 6, PC 7, or VR goggles 8. Hereinafter, the venue screen 5 will be mainly used as an example) (FIG. 3). In this case, the control unit 420 displays the display objects on the venue screen 5 at the display position in the captured image determined by the display position determination unit 417.

[0039] Furthermore, the term "athletes" may include, for example, multiple athletes. In this case, for example, the sound detection unit 413 detects sounds in the venue V based on sound information acquired from a vibration microphone, which is the microphone 3. Then, the association unit 418 associates the detected sound with one of the multiple athletes based on the sound detected by the vibration microphone and the detection result by the object detection unit 412. For example, if only one vibration microphone is installed in a position close to one of two athletes and the object detection unit 412 detects the sound of the athlete's shoes, the association unit 418 can associate the detected sound with one of the two athletes based on the sound (such as the volume) detected by the vibration microphone.

[0040] Furthermore, for example, the storage unit 42 may store a classification table that associates the detection results by the object detection unit 412 with the detection results by the sound detection unit 413. In this case, the display object determination unit 416 determines a display object that corresponds to the situation of the venue V based on the classification table.

[0041] Furthermore, for example, the association unit 418 associates actions with sounds based on the classification results by the object action classification unit 414 and the classification results by the sound classification unit 415. Then, the display position determination unit 417 determines the display position of the display object in the captured image based on the association results by the association unit 418.

[0042] Furthermore, for example, when displaying photographed images and displayed objects on the venue screen 5, the control unit 420 may display the displayed objects in a language corresponding to the player information of the athletes. For example, if an athlete from country A is playing against an athlete from country B, the displayed objects relating to the athlete from country A will be displayed in the language of country A, and the displayed objects relating to the athlete from country B will be displayed in the language of country B.

[0043] The display object may include, for example, at least one of characters, pictures, marks, and icons. The display object may also include, for example, characters corresponding to at least one of onomatopoeia, mimetic words, and mimetic sounds.

[0044] Furthermore, for example, the camera 2, microphone 3, and venue screen 5 are separate from one another.

[0045] Next, a table relating to the display forms of onomatopoeia and the processing of the display form determination unit 419 using the table will be described with reference to Fig. 4. Fig. 4 is a table relating to the display forms of onomatopoeia in this embodiment.

[0046] The display form determination unit 419 determines the display form of the display object based on, for example, at least one of the strength and length of the detected sound.

[0047] The display format includes, for example, the size, style (including font and language), color, display time, and animation (movement) of the characters, as shown in Fig. 4(a). Fig. 4(b) shows the final result (details will be described later).

[0048] First, we will explain the size of the onomatopoeia characters. If the volume of the detected sound is equal to or greater than the threshold, the size of the characters is set to 1.2 times the standard size. If the volume of the detected sound is less than the threshold, the size of the characters is set to 0.8 times the standard size.

[0049] Furthermore, if the posture determination result is a hitting sound (the sound of hitting with a racket), the character size is made 1.2 times the standard size, and if the posture determination result is a bouncing sound (the sound of bouncing on something other than a racket, such as a ping-pong table or the floor), the character size is made 0.8 times the standard size. Note that whether it is a "hitting sound" or a "bouncing sound" can be determined, for example, from the results of video analysis. Specifically, if the sound of a ball impact is detected and a racket swing is detected in the video, it is determined to be a "hitting sound," and if a racket swing is not detected in the video, it is determined to be a bouncing sound.

[0050] As a final result, the sum of the magnifications is calculated to determine the size of the characters. Therefore, for example, if the sound volume is equal to or greater than the threshold value and the posture determination result is the sound of a ball hitting the screen, the size of the characters will be larger.

[0051] Next, we will explain the style of onomatopoeia characters. If the volume of the detected sound is equal to or greater than the threshold, the character style is set to "sharp corners style." If the volume of the detected sound is less than the threshold, the character style is set to "rounded corners style."

[0052] If the detected sound level is equal to or higher than the threshold, the character style is set to "thin style," and if the detected sound level is lower than the threshold, the character style is set to "bold style."

[0053] In addition, if the posture determination result is the sound of a hitting ball, the character style is set to an "explosive design," and if the posture determination result is the sound of a bouncing ball, the character style is set to "no change."

[0054] As a final result, a style that includes each feature is selected.

[0055] Next, the color of the onomatopoeia characters will be explained. If the detected sound volume is equal to or greater than the threshold, the color of the characters will be red, and if the detected sound volume is less than the threshold, the color of the characters will be blue.

[0056] In addition, if the detected sound level is equal to or higher than the threshold, the transparency of the text is set to 30%, and if the detected sound level is lower than the threshold, the transparency of the text is set to 0%.

[0057] As a final result, the color and transparency are determined.

[0058] Next, we will explain the display time of onomatopoeia characters. If the detected sound volume is equal to or greater than the threshold, the display time of the characters is set to 1.2 times the standard display time, and if the detected sound volume is less than the threshold, the display time of the characters is set to 0.8 times the standard display time.

[0059] In addition, if the detected sound pitch is equal to or higher than the threshold, the display time of the characters is set to 1.2 times the standard display time, and if the detected sound pitch is lower than the threshold, the display time of the characters is set to 0.8 times the standard display time.

[0060] The final result is the sum of the magnifications, which determines the display time of the character. For example, if the loudness of the sound is above a threshold and the pitch of the sound is above a threshold, the display time of the character will be longer.

[0061] Next, the animation of the onomatopoeia characters will be described. If the volume of the detected sound is equal to or greater than the threshold, the animation of the characters is faded out, and if the volume of the detected sound is less than the threshold, the animation of the characters is turned off.

[0062] If the detected pitch of the sound is equal to or greater than the threshold, the animation of the characters is reduced, and if the detected pitch of the sound is less than the threshold, the animation of the characters is turned off.

[0063] In addition, if the posture determination result is the hitting sound, the animation of the characters is zoomed in, and if the posture determination result is the bouncing sound, the animation of the characters is turned off.

[0064] As a final result, a combination of each animation is determined.

[0065] 2, the control unit 420 superimposes and displays a display object on the captured image. In this case, the detection of the position and movement of the object by the object detection unit 412 and the superimposed display by the control unit 420 are performed in real time.

[0066] The display unit also includes a first display device (venue screen 5) installed in the venue V, and a second display device (smartphone 6, PC 7, VR goggles 8) installed in a remote location R different from the venue V. The control unit 420 and the second display device are connected via the Internet.

[0067] An input device for users to input information is connected to the second display device (smartphone 6, PC 7, VR goggles 8). The control unit 420 acquires the input result to the input device and displays the captured image on the first display device (venue screen 5) with the input result (for example, "Go for it!" displayed on display objects D1 and D4 in FIG. 3) superimposed thereon.

[0068] Furthermore, for example, the control unit 420 may move the display objects when displaying the captured image and the display objects on the venue screen 5. For example, the control unit 420 may move the display objects D1 to D4 in Fig. 3 while displaying them (for example, by zooming in, moving, vibrating, etc.).

[0069] Next, processing by the information processing device 4 will be described with reference to Figs. 5 to 8. Note that hereinafter, a target object may also be referred to as an "object." Fig. 5 is a flowchart showing processing by the information processing device 4 according to the embodiment. In step S1, the acquisition unit 411 acquires information on a plurality of externally connected devices (camera 2, microphone 3).

[0070] Next, in step S2, the acquisition unit 411 refers to the storage unit 42 and acquires setting information for the camera 2, microphone 3, and the like.

[0071] Next, in steps S3 to S15, the processing unit 41 executes each process while acquiring information from the camera 2, microphone 3, etc. in a loop process with a cycle of 100 ms (milliseconds).

[0072] In step S4, the processing unit 41 executes a recording process (a process related to the captured image acquired from the camera 2). For example, the object detection unit 412 detects an object based on the captured image.

[0073] Next, in step S5, the object motion classification unit 414 executes an object tracking process. Here, Fig. 6 is a flowchart showing the details of the process of step S5 in Fig. 5. In step S51, the object motion classification unit 414 uses the object detection model in the storage unit 42 to classify the object (racket, player, ping-pong table, ball, etc.), acquire the coordinate position, and track the object.

[0074] In parallel with step S5, in step S7, object motion classification unit 414 executes a posture determination process. Here, Fig. 7 is a flowchart showing the details of the process of step S7 in Fig. 5. In step S71, object motion classification unit 414 acquires the coordinate positions of the joints (points) of the object using the posture determination model in storage unit 42, and determines the posture of the object.

[0075] Furthermore, in step S8, the processing unit 41 executes a recording process (a process related to the sound information acquired from the microphone 3). For example, the sound detection unit 413 detects sounds in the venue based on the sound information.

[0076] Next, in step S9, the sound classification unit 415 executes sound inference processing (classification processing). Here, Fig. 8 is a flowchart showing the details of the processing in step S9 in Fig. 5. In step S91, the sound classification unit 415 vectorizes the sound information using the discrimination model of table tennis sounds in the storage unit 42.

[0077] Next, in steps S92 to S96, the sound classification unit 415 distinguishes between the sound of bouncing, the sound of shoes hitting the floor, the sound of applause, the sound of cheers, and the sound of a racket hitting the ball.

[0078] 5, after steps S7 and S9, in step S10, the display form determination unit 419 determines the display form of the onomatopoeia. Specifically, for example, the display form determination unit 419 uses the posture determination result and the sound inference result to determine the display form, such as the size, style, color, display time, and animation of the onomatopoeia, based on the table in FIG.

[0079] After steps S6 and S10, in step S11, control unit 420 executes a process for generating a video with onomatopoeia.

[0080] Next, in step S12, the control unit 420 performs a process of transmitting the streaming data to the video distribution site. In response to this, a user of the smartphone 6 or the VR goggles 8 of the PC 7 in the remote location R can access the video distribution site and watch the video with the onomatopoeia.

[0081] Next, in step S13, the control unit 420 performs display processing on the venue screen 5. As a result, spectators at venue V can watch the video with onomatopoeia on the venue screen 5. Also, for example, if the venue screen 5 is shown on television, television viewers can watch the video with onomatopoeia on television.

[0082] Next, in step S14, control unit 420 determines whether or not to end the process, and if Yes, ends the process, and if No, returns to step S3.

[0083] In this way, according to the video display system S of this embodiment, by using both the photographed images and sound information acquired at the venue V for a competition held at the venue V, it is possible to display objects corresponding to the actual venue conditions in relation to the photographed images.

[0084] Specifically, by classifying the sound, classifying the object's movement, and tracking the position of the moving object, it is possible to determine what kind of sound is at what position in the video, and therefore it is possible to display an object that corresponds to the actual situation at the venue in the captured image.

[0085] Furthermore, by detecting the position and movement of the object based on the captured image, the display position of the displayed object can be determined more appropriately.

[0086] Furthermore, by applying this technology to competition venues, it is possible to effectively communicate the situation and atmosphere of the competition to people with hearing impairments in particular.

[0087] Furthermore, by applying this to ball sports such as table tennis, it is possible to display display objects relating to the players, balls, and spectators as objects.

[0088] Furthermore, a more appropriate display object can be determined based on the classification results of the object's movement and the sound.

[0089] Furthermore, when there are multiple athletes, by using the sound detected by a vibration microphone, it is possible to use not only sound but also vibration information to identify with high accuracy which athlete is the source of the sound.

[0090] In addition, depending on the type of sound, those that can be classified using existing learning models are classified using the existing learning models, and those that cannot be classified using the transfer learning model are classified using the transfer learning model, thereby improving classification accuracy and reducing development costs.

[0091] Furthermore, by using a classification table that associates the results of object detection with the results of sound detection, it is possible to determine a display object that corresponds to the situation in the venue more quickly and with higher accuracy.

[0092] Furthermore, by associating the action with the sound based on the classification results of the object's action and the sound, the display position of the displayed object can be determined with higher accuracy.

[0093] Furthermore, if the display is displayed in a language that corresponds to the athlete's player information, the display can be made easier to understand for the athlete and spectators from the same country.

[0094] Also, by selecting from a variety of display items such as letters, pictures, marks, and icons, spectators can enjoy the game even more.

[0095] Moreover, by selecting the display items from onomatopoeia, mimetic words, and mimetic sounds that are likely to attract attention, spectators can be more entertained.

[0096] Furthermore, by providing the camera 2, microphone 3, and venue screen 5 as separate units, a flexible system design is possible.

[0097] Furthermore, by using a variety of display formats for the displayed objects, such as font size, style, color, display time, and animation, the spectators can be more entertained. Specifically, as explained using Figure 4, by varying the display format (size, style, color, display time, animation, etc.) of the onomatopoeia depending on the volume and pitch of the detected sound, the posture determination result, etc., the display format of the displayed objects can be flexibly changed according to the situation at the actual venue V.

[0098] Furthermore, by determining the display form of the display object depending on the strength and length of the detected sound, it is possible to display a more appropriate display object.

[0099] Furthermore, by detecting the position and movement of an object and superimposing the display object on the captured image in real time, this embodiment can be applied to live video as well as recorded video.

[0100] Furthermore, videos with onomatopoeia can be enjoyed by users in remote locations R using smartphones 6, PCs 7, and VR goggles 8, in addition to being viewed on the venue screen 5. In particular, if you use VR goggles 8 to watch videos with onomatopoeia in the metaverse space, you can enjoy a more immersive experience.

[0101] In addition, by reflecting the information entered on the smartphone 6, PC 7, or VR goggles 8 regarding cheering (for example, "Go for it!") as a display on the display screen, the interest of users of the smartphone 6, PC 7, or VR goggles 8 is increased.

[0102] Moreover, by displaying the displayed objects while moving, the images can be made more entertaining.

[0103] More specific display examples will be explained below. For example, if a player on one side of the table makes a "squeak" sound from his / her shoes when stepping forward to hit a shot to return the ball, it can be determined that the sound is coming from the shoes of that player, and the onomatopoeic display of "squeak" can be displayed near the shoes on the display screen.

[0104] Also, if a player on one side of the table makes a "kat" sound when the racket and ball collide when serving, the system can determine that the sound came from that player's racket and ball, and display the onomatopoeia "kat" near the racket and ball on the display screen.

[0105] Also, if a "clap clap" sound is heard from the spectators when one of the players makes a successful smash, the system can determine that the sound came from the spectators' seats and display the onomatopoeia "clap clap" near the spectators' seats on the display screen.

[0106] The program executed by the information processing device 4 of this embodiment can be provided by being recorded in an installable or executable file format on a computer-readable recording medium such as a CD (Compact Disc)-ROM (Read Only Memory), a flexible disk (FD), a CD-R (Recordable), or a DVD (Digital Versatile Disk).The program may also be provided or distributed via a network such as the Internet.

[0107] Although an embodiment of the present invention has been described above, this embodiment is presented as an example and is not intended to limit the scope of the invention. This novel embodiment can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. This embodiment and its modifications are included within the scope and spirit of the invention, and are also included in the invention and its equivalents as defined in the claims.

[0108] For example, although a venue for table tennis has been used as an example of a competition venue, the venue is not limited to this and may be a venue for other sports.Furthermore, the venue is not limited to a competition venue and may be another type of venue.

[0109] [Summary of this embodiment] The video display system of this embodiment comprises a photographing unit that photographs the venue, an object detection unit that detects a specified object based on the photographed image acquired from the photographing unit, a sound collection unit that collects sounds from the venue, a sound detection unit that detects sounds from the venue based on sound information acquired from the sound collection unit, a display object determination unit that determines a display object corresponding to the situation of the venue based on the detection results by the object detection unit and the detection results by the sound detection unit, a display unit that displays information, and a display control unit that displays the photographed image and the display object on the display unit.

[0110] According to this configuration, by using both the photographed images and sound information acquired at the venue for a competition held at the venue, it is possible to display an object corresponding to the actual situation at the venue in relation to the photographed images.

[0111] Furthermore, the venue is a competition venue where athletes play ball games in front of spectators, and the object includes at least one of the athletes, the ball, and the spectators. The object detection unit detects the position and movement of the object based on the captured image, and further includes a display position determination unit that determines the display position of the display object based on the detection result by the object detection unit. When the display control unit displays the captured image and the display object on the display unit, it displays the display object at the display position in the captured image. This is even more preferable.

[0112] This configuration provides the following additional benefits. First, by detecting the position and movement of the object based on the captured image, it is possible to more appropriately determine the display position of the displayed object. Furthermore, by applying this to a competition venue, it is possible to effectively convey the situation and atmosphere of the competition, especially to those with hearing impairments. Furthermore, by applying this to a ball sport such as table tennis, it is possible to display displayed objects related to the players, ball, and spectators as objects.

[0113] The system further includes an object motion classifier that classifies the motion of the object based on the detection result by the object detection unit, and a sound classifier that classifies detected sounds as sounds belonging to the venue or sounds belonging to the object based on the detection result by the sound detection unit. The display object determination unit determines a display object corresponding to the situation of the venue based on the classification result by the object motion classifier and the classification result by the sound classifier. This is even more preferable.

[0114] This configuration has the additional effect of being able to determine a more appropriate display object based on the classification results of the object's motion and the sound.

[0115] The athlete may include a plurality of athletes. The sound detection unit detects sounds in the venue based on sound information acquired from a vibration microphone serving as the sound collection unit. The display position determination unit associates the detected sound with one of the plurality of athletes based on the sound detected by the vibration microphone and the detection result by the object detection unit. This is even more preferable.

[0116] This configuration has the additional effect that when there are multiple athletes, by using the sound detected by the vibration microphone, it is possible to use not only sound but also vibration information to identify with high accuracy which athlete is the source of the sound.

[0117] Furthermore, it is even more preferable if the sound classification unit classifies the detected sounds based on a transfer learning model generated by transfer learning using a predetermined learning model.

[0118] With this configuration, depending on the type of sound, those that can be classified using existing learning models are classified using the existing learning models, and those that cannot be classified using existing learning models are classified using the transfer learning model, thereby achieving the additional effect of improving classification accuracy and reducing development costs.

[0119] The apparatus further includes a storage unit that stores a classification table that associates the detection results of the object detection unit with the detection results of the sound detection unit, and the display object determination unit determines the display object that corresponds to the situation of the venue based on the classification table.

[0120] This configuration has the additional effect of enabling the display object corresponding to the venue situation to be determined more quickly and accurately by using a classification table that associates the object detection results with the sound detection results.

[0121] The display position determination unit determines a display position of the object in the captured image based on the classification results of the object motion classification unit and the sound classification unit.

[0122] This configuration has the additional effect of associating actions and sounds based on the classification results of the object's actions and the classification results of the sounds, thereby making it possible to determine the display position of displayed objects with higher accuracy.

[0123] Furthermore, it is even more preferable if the display control unit, when displaying the captured image and the display object on the display unit, displays the display object in a language corresponding to the athlete information of the athlete.

[0124] This configuration has the additional effect of making the display easier to understand for the athlete and spectators from the same country by displaying the display in a language that corresponds to the athlete's player information.

[0125] It is even more preferable that the display includes at least one of characters, pictures, marks, and icons.

[0126] This configuration has the added effect of allowing spectators to enjoy the display even more by selecting from a wide variety of display items, such as characters, pictures, marks, and icons.

[0127] It is even more preferable if the display includes characters corresponding to at least one of onomatopoeia, mimetic words, and onomatopoeic words.

[0128] According to this configuration, by selecting the display objects from onomatopoeia, mimetic words, and mimetic sound words that are likely to attract attention, an additional effect is achieved in that spectators can be more entertained.

[0129] It is even more preferable if the sound collecting unit, the photographing unit, and the display unit are separate from one another.

[0130] This configuration has the additional advantage of enabling flexible system design by providing separate units for the imaging unit (camera 2), sound collection unit (microphone 3), and display unit (venue screen 5).

[0131] The display device may further include a display mode determining unit for determining a display mode of the displayed object, the display mode including a character size, style, color, display time, and animation.

[0132] This configuration has the added benefit of providing more enjoyment to spectators by using a variety of display modes for the displayed objects, such as character size, style, color, display time, and animation. Specifically, as explained with reference to Fig. 4, by varying the display mode (size, style, color, display time, animation, etc.) of the onomatopoeia depending on the volume and pitch of the detected sound, the posture determination result, etc., the display mode of the displayed objects can be flexibly changed depending on the situation at the actual venue V.

[0133] It is even more preferable if the display form determination unit determines the display form of the display object based on at least one of the intensity and the length of the detected sound.

[0134] This configuration has the additional effect of determining the display form of the display object depending on the strength and length of the detected sound, thereby making it possible to display a more appropriate display object.

[0135] Furthermore, the display control unit displays the display object superimposed on the captured image. The detection of the position and movement of the object by the object detection unit and the superimposed display by the display control unit are performed in real time. This is even more preferable.

[0136] According to this configuration, by detecting the position and movement of the object and superimposing the display object on the captured image in real time processing, an additional effect is achieved in that this embodiment can be applied not only to recorded video but also to live video.

[0137] Furthermore, it is even more preferable that the display unit includes a first display device installed in the venue and a second display device installed in a remote location different from the venue, and the display control unit and the second display device are connected via the Internet.

[0138] This configuration provides the following additional benefits: By making videos with onomatopoeia viewable not only on a first display device installed at the venue, but also on a second display device (smartphone 6, PC 7, VR goggles 8) installed in a remote location different from the venue, it is possible to entertain many people. In particular, by using VR goggles 8 to watch videos with onomatopoeia in the metaverse space, a greater sense of immersion can be enjoyed.

[0139] Furthermore, an input device for a user to input information is connected to the second display device, and the display control unit acquires the input result to the input device and displays the captured image on the first display device while superimposing the input result.

[0140] According to this configuration, the information input about cheering (for example, "Go for it!") on the second display device (smartphone 6, PC 7, or VR goggles 8) is reflected as a display on the display screen, thereby providing the additional effect of increasing the interest of users of the smartphone 6, PC 7, or VR goggles 8.

[0141] Furthermore, it is even more preferable if the display control unit, when displaying the captured image and the display object on the display unit, displays the display object while moving it.

[0142] This configuration has the additional effect of making the image more entertaining by displaying the display object while moving it. [Explanation of symbols]

[0143] 2...camera (photographing device), 3...microphone (sound collection unit), 4...information processing device (video display device), 5...venue screen (first display device), 6...smartphone (second display device), 7...PC (second display device), 8...VR goggles (second display device), 42...memory unit, 44...display unit, 412...object detection unit, 413...sound detection unit, 414...object movement classification unit, 415...sound classification unit, 416...display object determination unit, 417...display position determination unit, 418...association unit, 419...display form determination unit, 420...control unit (display control unit), R...remote location, S...video display system, V...venue (competition venue)

Claims

1. The photography team will be taking photos of the venue, an object detection unit that detects a predetermined object based on the captured image acquired from the imaging unit; a sound collection unit that collects sounds from the venue; a sound detection unit that detects sounds in the venue based on the sound information acquired from the sound collection unit; a display object determination unit that determines a display object corresponding to the situation of the venue based on the detection result by the object detection unit and the detection result by the sound detection unit; a display unit that displays information; a display control unit that displays the captured image and the display object on the display unit;

2. the venue is a competition venue where a ball game is played by players in front of spectators, the object includes at least one of the player, the ball, and the spectator; the object detection unit detects a position and a movement of the object based on the captured image; a display position determination unit that determines a display position of the display object based on a detection result by the object detection unit; The video display system according to claim 1 , wherein the display control unit, when displaying the photographed image and the display object on the display unit, displays the display object at the display position in the photographed image.

3. The competitors include a plurality of competitors; the sound detection unit detects sounds in the venue based on sound information acquired from a vibration microphone serving as the sound collection unit; 3. The video display system according to claim 2, wherein the display position determination unit associates the detected sound with one of the plurality of competitors based on the sound detected by the vibration microphone and the detection result by the object detection unit.

4. 2. The video display system of claim 1, wherein the display unit includes a first display device installed at the venue and a second display device installed at a remote location different from the venue, and the display control unit and the second display device are connected via the Internet.

5. an object detection unit that detects a predetermined object based on a captured image acquired from a capture unit that captures images of the venue; a sound detection unit that detects sounds from the venue based on sound information acquired from a sound collection unit that collects sounds from the venue; a display object determination unit that determines a display object corresponding to the situation of the venue based on the detection result by the object detection unit and the detection result by the sound detection unit; A video display device comprising: a display unit that displays information; and a display control unit that displays the captured image and the displayed object.

6. an object detection step in which the object detection unit detects a predetermined object based on a captured image acquired from the image capturing unit that captures an image of the venue; a sound detection step in which a sound detection unit detects sounds from the venue based on sound information acquired from a sound collection unit that collects sounds from the venue; a display object determination step in which a display object determination unit determines a display object corresponding to the situation of the venue based on the detection result by the object detection unit and the detection result by the sound detection unit; a display control step in which a display control unit displays the captured image and the display object on a display unit that displays information.

Citation Information

Patent Citations

  • Apparatus and program for image decoration

    JP2011205296A

  • Onomatopoeia presentation device, onomatopoeia presentation program, and onomatopoeia presentation method relating to surrounding environment evaluation results

    JP6917311B2