Switching device, switching system, and program
The described switching device automatically selects video images based on visual and audio cues from multiple performers, addressing the resource-intensive and costly limitations of existing techniques by enabling efficient and automated video switching.
Patent Information
- Application Number
- JP2023192956
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video signal switching techniques require extensive human resources and are costly, as they necessitate the preparation of multiple robot cameras and complex control rules.
A switching device that automatically selects and switches video images based on both visual and audio cues from multiple performers, using an acquisition unit, a switching control unit, and an output section to generate and output switching signals.
Enables automatic and efficient video switching, reducing the need for extensive human resources and minimizing costs, while maintaining focus on the performers and the program content.
Smart Images

Figure 2025080012000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a video signal switching technique. [Background technology]
[0002] Conventionally, in the field of program production at a broadcasting station, multiple cameras were used to shoot multiple performers, and a switcher independently selected one image from the multiple camera images to output it. For this reason, a large number of human resources were required in the field of program production as described above. From this point of view, Patent Document 1 proposes a method of controlling the shooting shots of a subject shot using multiple robot cameras. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2008-72702 A Summary of the Invention [Problem to be solved by the invention]
[0004] However, the invention of Patent Document 1 requires the preparation of many robot cameras, which is costly, and also requires detailed definition of rules for controlling each robot camera.
[0005] An object of the present invention is to provide a switching device capable of automatically switching images based on the images and sounds of performers. [Means for solving the problem]
[0006] In one aspect of the invention, a switching device comprises: An acquisition unit that acquires individual images of a plurality of people and audio of the plurality of people; a switching control unit that generates a switching signal that selects one image from a plurality of images including the individual images of the plurality of people based on at least one of the individual images and the audio; and an output section that outputs the switching signal.
[0007] In another aspect of the invention, a program includes: Acquiring individual images of a plurality of persons and audio of the plurality of persons; generating a switching signal for selecting one image from a plurality of images including the individual images of the plurality of people based on at least one of the individual images and the audio; The computer is caused to execute a process for outputting the switching signal. Effect of the Invention
[0008] According to the present invention, it is possible to automatically switch between videos based on the video and audio of the performers. [Brief description of the drawings]
[0009] [Figure 1] 1 shows a schematic configuration of a switching system according to a first embodiment. [Diagram 2] 1 shows an internal configuration of an automatic switching device according to a first embodiment. [Diagram 3] 1 shows an example of voice-based switching. [Figure 4] 1 shows another example of voice-based switching. [Diagram 5] 1 shows another example of voice-based switching. [Figure 6] 1 shows an example of reaction-based switching. [Figure 7] 13 shows another example of reaction-based switching. [Figure 8] 13 shows another example of reaction-based switching. [Figure 9] 13 is a flowchart of a switching process. [Figure 10]An example of a flowchart of the switching process according to the second embodiment. [Figure 11] Shows the internal configuration of the automatic switching device according to the third embodiment. [Figure 12] Schematically shows how the performer operates the flip. [Figure 13] Shows the internal configuration of the automatic switching device according to the fourth embodiment.
Embodiments for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. [First Embodiment] (Configuration of the Switching System) FIG. 1 shows a schematic configuration of a switching system according to the first embodiment. The switching system 1 is a system installed at a TV program production site or the like, which produces a program video by appropriately switching the videos of the program performers. As shown in the figure, the switching system 1 includes an automatic switching device 100 and a switcher 5.
[0011] In the example of FIG. 1, there are person A and person B as performers, and videos are taken by 3 cameras. The A camera is a camera that takes pictures of person A, the B camera is a camera that takes pictures of person B, and the 2-shot (hereinafter referred to as "2S") camera is a camera that takes 2S videos including person A and person B at the same time. The 2S video is an example of the overall video of the present invention. The video (individual video) Va taken by the A camera, the video (individual video) Vb taken by the B camera, and the 2S video Vab taken by the 2S camera are input to the automatic switching device 100 and the switcher 5. Also, the voice Aa of person A and the voice Ab of person B collected by the microphone are input to the automatic switching device 100.
[0012] The automatic switching device 100 generates a switching command SC for selecting one of the images Va, Vb, Vab based on the input images Va, Vb, Vab and audio Aa, Ab, and outputs the generated switching command to the switcher 5. The switcher 5 selects one of the images Va, Vb, Vab in accordance with the switching command SC, and outputs the selected image to the outside. The switching command is an example of a switching signal of the present invention.
[0013] (Configuration of automatic switching device) 2 is a block diagram showing the internal configuration of the automatic switching device 100. As shown in the figure, the automatic switching device 100 includes a voice analysis unit 11, a feature point analysis unit 12, a facial expression analysis unit 13, and a switching control unit 14. The switching control unit 14 selects one of the input images Va, Vb, Vab based on at least one analysis result of the voice analysis unit 11, the feature point analysis unit 12, and the facial expression analysis unit 13.
[0014] The voice analysis unit 11 receives the voice Aa of person A and the voice Ab of person B. The voice analysis unit 11 compares the voice levels of the input voices Aa and Ab with a predetermined voice level (hereinafter also referred to as the "reference level") and outputs the comparison result to the switching control unit 14. Here, the reference level is set to a standard voice level during normal conversation, for example. That is, the voice analysis unit 11 determines that a person corresponding to a voice exceeding the reference level is talking (hereinafter also referred to as "talking"), and outputs information indicating the person who is talking to the switching control unit 14. The switching control unit 14 generates a switching command SC for selecting one of the images Va, Vb, and Vab based on the information input from the voice analysis unit 11. In this way, switching performed based on the analysis result of the voice is also called "voice-based switching". Then, the switching control unit 14 outputs the generated switching command SC to the switcher 5. In this way, the switching control unit 14 can select the image of the person who is talking among multiple people.
[0015] FIG. 3 shows an example of audio-based switching. Specifically, when only audio Aa exceeds the reference level, the switching control unit 14 determines that only person A is talking, and selects the video Va corresponding to person A as shown in FIG. 3(A). For convenience of illustration in the drawings attached to this specification, a person wearing a mask indicates a person who is not talking. Also, when only audio Ab exceeds the reference level, the switching control unit 14 determines that only person B is talking, and selects the video Vb corresponding to person B as shown in FIG. 3(B). On the other hand, when both audio Aa and audio Ab exceed the reference level, the switching control unit 14 determines that both person A and person B are talking, and selects the 2S video Vab corresponding to person A and person B as shown in FIG. 3(C).
[0016] The switching control unit 14 generates a switching signal including an instruction to select one of the multiple images and an instruction to maintain the selection of the image for a first predetermined time t1 as a switching command SC. As a result, the switcher 5 selects one image based on the switching command SC, and then cancels the selection after the predetermined time t1 has elapsed. In this embodiment, if neither the image Va of person A nor the image Vb of person B is selected by the switching command SC, the switcher 5 selects the 2S image Vab including both person A and person B as the default image. Therefore, when the switcher 5 switches to the image Va of person A at a certain timing, it maintains the selection of the image Va for the predetermined time t1, and then switches to the 2S image Vab. In this way, by maintaining the state for the predetermined time t1 after the image is switched once, it is possible to prevent switching from occurring frequently and hastily switching the program images. In addition, as the default image, another image prepared in advance may be used instead of the 2S image Vab including person A and person B.
[0017] In the above method, if a certain person continues to talk, only the image of that person will be selected for that entire time. Therefore, the switching control unit 14 may switch to the 2S image Vab when a second predetermined time t2 (e.g., 10 seconds) has elapsed after selecting one image based on the voice-based switching. For example, as shown in FIG. 4, if only person A continues to talk for a long time, the switching control unit 14 may automatically switch to the 2S image Vab after the predetermined time t2 has elapsed. This makes it possible to prevent the image of the same person from being displayed for a long time.
[0018] During the production of a program, various kinds of noise may occur in addition to the talk of the people appearing on the program. For example, there is paper noise when a person turns the pages of a paper document, and operation sounds (keystroke sounds) when operating a keyboard. Therefore, an audio filter that cuts the frequencies of such noise sounds may be provided in the audio analysis unit 11 to remove the noise sounds from the input audio Aa and Ab. In this way, as shown in FIG. 5(A), even if the keystroke sounds of person A are loud, the audio analysis unit 11 does not detect the keystroke sounds, and the switching control unit 14 can correctly select the video of person B who is talking.
[0019] By the above-mentioned voice-based switching, the switching control unit 14 basically selects the image of the person who is talking, but a slight "pause" may occur while the person is talking. For example, this may occur when the person who is talking is thinking about what to say next or when choosing words. However, if a person makes a short utterance such as "Oh," during a "pause" while one person is talking, the switching control unit 14 reacts to this utterance, and the image is frequently switched, making it difficult to see. Therefore, even if the voice level falls below the reference level during a "pause" during the talk after a person starts talking and the voice level exceeds a reference level, the switching control unit 14 sets a third predetermined time t3 (for example, 0.5 seconds) as a non-response period. In other words, even if the voice of another person exceeds the reference level during the non-response period, the switching control unit 14 does not switch in response to it. This allows the switching control unit 14 to not react to the short utterance of another person during a short "pause" while one person is talking. For example, as shown in FIG. 5(B), even if person A responds with "Huh?" during a pause while person B is talking, the image will not switch to that of person A in response.
[0020] Returning to FIG. 2, the feature point analysis unit 12 receives the video Va of person A, the video Vb of person B, and the 2S video Vab. The feature point analysis unit 12 is an example of a reaction detection unit, and performs video analysis of the video Va of person A and the video Vb of person B to detect a predetermined state of the person (also called a "predetermined reaction"). Specifically, the feature point analysis unit 12 analyzes the input video to extract facial feature point coordinates, and detects a state such as a surprised state, a state about to speak, or a nodding state of the person from the positional relationship and movement of the feature point coordinates, and outputs information indicating the state to the switching control unit 14. When a predetermined state is detected, the switching control unit 14 selects a video of the person in that state. This switching is also called "reaction-based switching".
[0021] As an example, the feature point analysis unit 12 extracts feature points corresponding to the upper lip and the lower lip from the image of the person's face, and detects the degree of opening of the mouth from the distance between the upper lip and the lower lip. Then, if the degree of opening of the mouth exceeds a predetermined value, it is determined that the person is in a surprised state or a state of about to speak, and outputs information indicating that state to the switching control unit 14. The switching control unit 14 selects the image of the person in a surprised state or a state of about to speak. FIG. 6 shows an example of reaction-based switching. In the example on the left side of FIG. 6, person B is talking, and the switching control unit 14 selects the image of person B. Thereafter, as shown on the right side of FIG. 6, when person A is in a surprised state with his mouth open, the switching control unit 14 selects the image of person A.
[0022] As another example, the feature point analysis unit 12 extracts a feature point corresponding to the tip of the nose from the image of the person's face, and detects the movement of the tip of the nose. Then, when the number of times the tip of the nose moves exceeds a specified value, it determines that the person is in a nodding state, and outputs information indicating that state to the switching control unit 14. The switching control unit 14 selects the image of the person in a nodding state. FIG. 7 shows another example of reaction-based switching. In the example on the left side of FIG. 7, person A is talking, and the switching control unit 14 selects the image of person A. Thereafter, when person B is in a nodding state as shown on the right side of FIG. 7, the switching control unit 14 selects the image of person B.
[0023] Returning to FIG. 2, the facial expression analysis unit 13 receives the video Va of person A, the video Vb of person B, and the 2S video Vab. The facial expression analysis unit 13 is an example of a reaction detection unit, and performs video analysis of the video Va of person A and the video Vb of person B to detect a predetermined state (predetermined reaction) of the person. Specifically, the facial expression analysis unit 13 detects the smile of the person using AI (Artificial Intelligence) face recognition software or the like. Since the AI face recognition software can output a smile index that scores a smile from a face image, the facial expression analysis unit 13 calculates the smile index of each person based on the image of each person. Then, when the smile index exceeds a specified value, the facial expression analysis unit 13 determines that the person is in a smiling state, and outputs information indicating that to the switching control unit 14. The switching control unit 14 selects the image of the person in a smiling state.
[0024] Fig. 8 shows another example of reaction-based switching. In the example on the left side of Fig. 8, the smile index of person A is 10 points and the smile index of person B is 20 points, neither of which exceeds a specified value (e.g., 80 points), and since person B is talking, the video of person B is selected. Thereafter, as shown in the example on the right side of Fig. 8, when the smile index of person A exceeds the specified value, the facial expression analysis unit 13 determines that person A is in a smiling state and outputs the information to the switching control unit 14. The switching control unit 14 selects the video of person A based on the information that person A is in a smiling state.
[0025] (Switching process) 9 is a flowchart of the switching process by the automatic switching device 100. This process is realized by a computer such as a CPU executing a program prepared in advance and operating as each element shown in FIG.
[0026] First, the automatic switching device 100 acquires video and audio of multiple people (step S11). Next, the automatic switching device 100 analyzes the audio and reaction based on the audio or video using the audio analysis unit 11, the feature point analysis unit 12, and the facial expression analysis unit 13 (step S12). Then, the automatic switching device 100 determines whether or not a switching condition is met based on at least one of the analysis results of the audio analysis unit 11, the feature point analysis unit 12, and the facial expression analysis unit 13 (step S13). Here, the switching condition is a condition for the switching control unit 14 to switch the video, and specifically includes that the audio level of a certain person exceeds a reference level based on the audio analysis, and that the reaction of a certain person corresponds to a predetermined reaction based on the feature point analysis or the facial expression analysis.
[0027] When the switching condition is met (step S14: Yes), the switching control unit 14 generates a switching command SC for selecting the video of the person who meets the switching condition, and outputs it to the switcher 5 (step S14). Then, the automatic switching device 100 judges whether an instruction to end the process has been input (step S15). The instruction to end the process is input by an operator, for example, when recording of a program ends. If the instruction to end the process has not been input (step S15: No), the process returns to step S11, and steps S11 to S14 are repeated. On the other hand, if an instruction to end the process has been input (step S15: Yes), the switching process ends.
[0028] [Second embodiment] In the first embodiment, voice-based switching and reaction-based switching are used together, and if either of them occurs, the video is switched based on that. However, if the video is switched in response to all detection results, problems such as hasty switching, poor speaker selection, and a lack of focus on the story may occur. Therefore, in the second embodiment, voice-based switching and reaction-based switching are combined according to a predetermined rule to perform more appropriate switching.
[0029] Specifically, in the second embodiment, switching is performed according to the following rules. (Rule 1) When an individual video of a certain person is switched to by voice-based switching, reaction-based switching is enabled. However, reaction-based switching is not enabled for a fourth predetermined time t4 (e.g., 5 seconds) from the timing of switching to the individual video of a certain person. In other words, reaction-based switching is enabled after the predetermined time t4 has elapsed since switching to the individual video of a certain person by voice-based switching. (Rule 2) The reaction-based switching interval is, for example, 5 seconds. After switching is performed by reaction-based switching, the state is not continued for a long time, and the selected state is released after a fifth predetermined time t5 (for example, 2 seconds) has elapsed. By combining voice-based switching and reaction-based switching according to the above rules, natural switching is possible.
[0030] FIG. 10 shows an example of switching processing when applying the above-mentioned rules 1 and 2 and combining voice-based switching and reaction-based switching. This processing is realized by a computer such as a CPU executing a program prepared in advance and operating as each element shown in FIG. 2. It is assumed that two people are appearing in the program. Also, the switching control unit 14 uses 2S video as the default video. That is, if neither the voice-based nor reaction-based switching conditions are met, the switching control unit 14 selects the 2S camera.
[0031] First, the voice analysis unit 11 determines whether or not both of the people appearing are talking (step S21). If neither of them are talking (step S21: No), the process proceeds to step S24. On the other hand, if both of them are talking (step S21: Yes), the switching control unit 14 determines whether or not this state has continued for 1.5 seconds (step S22). If this state has continued for 1.5 seconds (step S22: Yes), the switching control unit 14 selects the 2S camera, keeps this state for 4.5 seconds or more, and then proceeds to step S32.
[0032] In step S24, the voice analysis unit 11 judges whether or not both of the performers are talking. If one of the two people is talking (step S24: No), the process proceeds to step S27. On the other hand, if neither of the two people is talking (step S24: Yes), the switching control unit 14 judges whether or not this state has continued for two seconds (step S25). If this state has continued for two seconds (step S25: Yes), the switching control unit 14 selects the 2S camera, keeps this state for 4.5 seconds or more, and then proceeds to step S32.
[0033] In step S27, since one of the two people is talking, the switching control unit 14 selects the camera of the person who is talking and keeps it for 5 seconds (step S27). Next, the feature point analysis unit 12 and the facial expression analysis unit 13 determine whether the other person has made a predetermined reaction (step S28). Here, in accordance with the above-mentioned rule 1, when an individual image of a certain person is switched to by voice-based switching, the reaction detection is performed. Also, in accordance with the rule 1, the reaction detection is enabled 5 seconds after the timing of switching to the individual image. Furthermore, in accordance with the rule 1, the reaction detection interval is set to 5 seconds. The predetermined reaction may be at least one of the above-mentioned surprised state, about-to-speak state, nodding state, and smiling state.
[0034] If the other person reacts (step S28: Yes), the switching control unit 14 selects the camera of the person who reacted, keeps that state for two seconds, and then cancels the selection of that camera (step S29). Here, in accordance with the above-mentioned rule 2, after switching by reaction-based switching, the selected state is canceled two seconds later. Then, the process proceeds to step S32.
[0035] If it is determined in step S28 that the other person has not reacted (step S28: No), the switching control unit 14 determines whether or not that state, i.e., the state in which one person is talking, has continued for 10 seconds or more (step S31). If that state has continued for 10 seconds or more (step S31: Yes), in order to avoid the same person's video being selected for a long time, the switching control unit 14 selects the 2S camera, keeps that state for 4.5 seconds or more, and then proceeds to step S32.
[0036] In step S32, the automatic switching device 100 judges whether an instruction to end the process has been input (step S33). The instruction to end the process is input by an operator, for example, at the end of recording a program. If an instruction to end the process has not been input (step S32: No), the process returns to step S21. On the other hand, if an instruction to end the process has been input (step S32: Yes), the process ends.
[0037] As described above, in the second embodiment, more appropriate switching can be performed by combining voice-based switching and reaction-based switching in accordance with a predetermined rule.
[0038] [Third embodiment] In the above first and second embodiments, the video of people such as performers in a program is switched, but automatic switching to a VTR that is played during the program cannot be performed. Therefore, in the third embodiment, automatic switching to a pre-prepared VTR is performed.
[0039] Fig. 11 shows the internal configuration of an automatic switching device 100a according to the third embodiment. As can be understood by comparing with Fig. 2, the automatic switching device 100a of the third embodiment is similar to the automatic switching device 100 of the first embodiment, except that it includes a specific word detection unit 16. Therefore, a description of the same points as in the first embodiment will be omitted. In the third embodiment, video from a prepared VTR is input to a switcher 5.
[0040] The voice Aa of person A and the voice Ab of person B are input to the specific word detection unit 16. The specific word detection unit 16 performs real-time transcription based on the voices Aa and Ab, and converts the contents of the conversations of persons A and B into character strings. The specific word detection unit 16 then detects predetermined specific words from the obtained character strings. The specific words are words related to the playback of a VTR, and can be set to, for example, "VTR," "please take a look," etc.
[0041] The specific word detection unit 16 outputs the detection result of the specific word to the switching control unit 14. When the specific word is detected, the switching control unit 14 outputs a switching command SC to the switcher 5 to instruct switching to a previously prepared VTR. When the switching command SC to instruct switching to the VTR video is input, the switcher 5 selects and outputs the VTR video. In this way, when any of a plurality of people such as performers utters the specific word, the automatic switching device 100a automatically switches to the VTR video accordingly. In this way, automatic switching to the VTR video is possible.
[0042] [Fourth embodiment] As part of the direction of television programs, performers may use flip charts to give commentary or express their opinions. As a method of program production, when a performer holds up a flip chart, in order to make the flip chart easier for viewers to read, switching may be made to a camera that captures images close up to the flip chart (hereinafter referred to as a "flip camera"). The fourth embodiment performs automatic switching to images from a flip camera.
[0043] Fig. 12 shows a model of a performer operating a flip. In Fig. 12, the performer holds up a flip 21, and the flip camera captures a close-up image of the flip (also called "flip image") and outputs it to the switcher 5. The flip camera is a camera dedicated to capturing flip image.
[0044] As shown in the figure, a sensor terminal 22 is attached to the back surface of the flip 21. The sensor terminal 22 is a terminal equipped with a gyro sensor, and may be a small commercially available terminal, a smartphone or tablet on which an application capable of acquiring information from the gyro sensor is installed, or the like.
[0045] Fig. 13 shows the internal configuration of an automatic switching device 100b according to the fourth embodiment. As can be understood by comparison with Fig. 2, the automatic switching device 100a of the fourth embodiment is similar to the automatic switching device 100 of the first embodiment, except that it includes a tilt determination unit 17. Therefore, a description of the same points as in the first embodiment will be omitted. In the fourth embodiment, a flip image captured by a flip camera is input to the switcher 5.
[0046] The automatic switching device 100b is capable of wireless communication with the sensor terminal 22 attached to the back surface of the flip 21, and receives a sensor signal of the gyro sensor from the sensor terminal 22. The sensor signal indicates the tilt angle of the sensor terminal 22, i.e., the tilt angle of the flip 21. The tilt determination unit 17 determines that the flip 21 is upright when the tilt angle of the flip 21 is equal to or greater than a predetermined angle (e.g., 80°), and determines that the flip 21 is not upright when the tilt angle is less than the predetermined angle. The tilt determination unit 17 then outputs a determination result indicating whether or not the flip 21 is upright to the switching control unit 14.
[0047] When the flip 21 is upright, the switching control unit 14 generates a switching command SC instructing switching to the flip video from the flip camera, and outputs it to the switcher 5. When the flip 21 is no longer upright, the switching control unit 14 generates a switching command SC for canceling the selection of the flip video, and outputs it to the switcher 5. When the selection of the flip video is canceled, the switching control unit 14 returns to switching based on voice-based switching or reaction-based switching. Thus, according to the fourth embodiment, when a performer sets up the flip 21, it is possible to automatically switch to the flip video.
[0048] [Variations] Modifications of the above embodiment will be described below. The following modifications can be applied to the above embodiment in appropriate combination.
[0049] (Variation 1) In the automatic switching devices 100, 100a, and 100b of the above embodiments, switching may be performed using facial recognition of the performers. Specifically, the facial expression analysis unit 13 continuously performs facial recognition in the individual video of a person. When the switching control unit 14 selects an individual video of a certain person, if the person stands up or the like and the person's face goes out of the frame of the selected individual video, the facial expression analysis unit 13 detects that the person's face is no longer included in the selected video and notifies the switching control unit 14. In response to this, the switching control unit 14 automatically switches to another video that includes the person's face. This makes it possible to automatically switch to an appropriate video even if the face goes out of the frame due to the person's movement or the like.
[0050] (Variation 2) In the above embodiment, the people to be switched are the performers, but if footage of staff, spectators, etc. is to be inserted depending on the content of the program, a camera may be provided to film those people and the footage may be switched to that person. Even in this case, the above-mentioned voice-based switching and reaction-based switching techniques can basically be applied.
[0051] (Variation 3) In the above embodiment, it is assumed that there are two performers and 2S video is used as the overall video, but if there are more performers, a video including some of the performers may be used as the overall video, or a video including all of the performers may be used as the overall video.
[0052] In addition, some or all of the above-described embodiments (including modified examples, the same applies below) can be described as, but are not limited to, the following notes.
[0053] (Appendix 1) An acquisition unit that acquires individual images of a plurality of people and audio of the plurality of people; a switching control unit that generates a switching signal that selects one image from a plurality of images including the individual images of the plurality of people based on at least one of the individual images and the audio; an output unit that outputs the switching signal; A switching device comprising:
[0054] (Appendix 2) The switching device according to claim 1, wherein the switching control unit generates a switching signal including an instruction to switch to the one image and an instruction to maintain selection of the one image for a first predetermined time.
[0055] (Appendix 3) a voice analysis unit that analyzes the level of the voice, The switching device according to claim 1, wherein the switching control unit generates a switching signal that selects a video of a person whose audio level exceeds a predetermined reference level.
[0056] (Appendix 4) The switching device of claim 3, wherein when a person's voice level exceeds the reference level and then falls below the reference level, the switching control unit does not generate a switching signal based on analysis of the voice level of other people for a third predetermined time period.
[0057] (Appendix 5) The acquisition unit acquires an entire video including the plurality of people at the same time, The switching device according to claim 1, wherein the switching control unit selects the one image from among the individual images of the plurality of people and the overall image.
[0058] (Appendix 6) The switching control unit of the switching device described in Appendix 5 generates a switching signal including, when selecting an individual image of one of the multiple people as the one image, an instruction to select the individual image of the one person and an instruction to switch the individual image to the overall image after a second predetermined time has elapsed.
[0059] (Appendix 7) A reaction detection unit detects a predetermined reaction of the person by video analysis of the individual video, The switching control unit generates a switching signal that selects an image of a person corresponding to the predetermined reaction.
[0060] (Appendix 8) The switching device according to claim 7, wherein the reaction detection unit detects at least one of the person's surprised state, nodding state, and smiling state as the predetermined reaction.
[0061] (Appendix 9) a voice analysis unit for analyzing a level of the voice; A reaction detection unit that detects a predetermined reaction of the person by video analysis of the individual video, The switching control unit performs voice-based switching to select a video of a person whose voice level exceeds a predetermined reference level, and reaction-based switching to select a video of a person whose reaction corresponds to the predetermined reaction, The switching device according to claim 1, wherein the switching control unit performs the reaction-based switching when the one video is selected by the voice-based switching.
[0062] (Appendix 10) The switching device according to claim 9, wherein the switching control unit performs the reaction-based switching detection after a fourth predetermined time has elapsed after the one video is selected by the voice-based switching.
[0063] (Appendix 11) The switching device according to claim 9, wherein the switching control unit selects the one video by the reaction-based switching and then deselects the one video after a fifth predetermined time has elapsed.
[0064] (Appendix 12) a word detection unit for detecting a specific word from the speech, 2. The switching device according to claim 1, wherein the switching control unit generates a switching signal for selecting a pre-prepared VTR video when the specific word is detected.
[0065] (Appendix 13) a tilt determination unit that receives a sensor signal indicating a tilt of a flip operated by at least one of the plurality of persons from a sensor provided on the flip and determines whether or not the tilt of the flip has exceeded a predetermined angle; 2. The switching device according to claim 1, wherein the switching control unit generates a switching signal for selecting a flip image captured by shooting the flip when an inclination of the flip exceeds a predetermined angle.
[0066] (Appendix 14) A switching device according to claim 1; a switcher that receives the individual images of the plurality of persons and the switching signal and outputs the one image based on the switching signal; A switching system comprising:
[0067] (Appendix 15) Acquiring individual images of a plurality of persons and audio of the plurality of persons; generating a switching signal for selecting one image from a plurality of images including the individual images of the plurality of people based on at least one of the individual images and the audio; A program that causes a computer to execute a process of outputting the switching signal.
[0068] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-mentioned embodiments. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. In other words, the present invention naturally includes various modifications and corrections that a person skilled in the art could make in accordance with the entire disclosure, including the claims, and the technical ideas. In addition, the disclosures of the above cited patent documents and the like are incorporated herein by reference. [Explanation of symbols]
[0069] 1. Switching System 5. Switcher 11 Voice Analysis Section 12 Feature point analysis section 13 Facial expression analysis department 14 Switching control section 16 Specific word detection section 17 Tilt determination section 21 Flip 22 Sensor terminal 100, 100a, 100b Automatic switching device
Claims
1. An acquisition unit that acquires individual images of a plurality of people and audio of the plurality of people; a switching control unit that generates a switching signal that selects one image from a plurality of images including the individual images of the plurality of people based on at least one of the individual images and the audio; an output unit that outputs the switching signal; A switching device comprising:
2. The switching device according to claim 1 , wherein the switching control unit generates a switching signal including an instruction to switch to the one video and an instruction to maintain selection of the one video for a first predetermined time.
3. a voice analysis unit that analyzes the level of the voice, The switching device according to claim 1 , wherein the switching control unit generates a switching signal for selecting a video of a person whose audio level exceeds a predetermined reference level.
4. 4. The switching device according to claim 3, wherein when a voice level of a certain person exceeds the reference level and then falls below the reference level, the switching control unit does not generate a switching signal based on an analysis of the voice level of other people for a third predetermined time.
5. The acquisition unit acquires an entire video including the plurality of people at the same time, The switching device according to claim 1 , wherein the switching control unit selects the one image from among the individual images of the plurality of people and the overall image.
6. The switching device of claim 5, wherein the switching control unit, when selecting an individual image of one of the plurality of persons as the one image, generates a switching signal including an instruction to select the individual image of the one person and an instruction to switch the individual image to the overall image after a second predetermined time has elapsed.
7. A reaction detection unit detects a predetermined reaction of the person by video analysis of the individual video, The switching device according to claim 1 , wherein the switching control unit generates a switching signal for selecting an image of a person corresponding to the predetermined reaction.
8. The switching device according to claim 7 , wherein the reaction detection unit detects at least one of a surprised state, a nodding state, and a smiling state of the person as the predetermined reaction.
9. a voice analysis unit for analyzing a level of the voice; A reaction detection unit that detects a predetermined reaction of the person by video analysis of the individual video, The switching control unit performs voice-based switching to select a video of a person whose voice level exceeds a predetermined reference level, and reaction-based switching to select a video of a person whose reaction corresponds to the predetermined reaction, The switching device according to claim 1 , wherein the switching control unit performs the reaction-based switching when the one video is selected by the voice-based switching.
10. The switching device according to claim 9 , wherein the switching control unit performs the reaction-based switching after a fourth predetermined time has elapsed after the one video is selected by the voice-based switching.
11. The switching device according to claim 9 , wherein the switching control unit cancels the selection of the one video after a fifth predetermined time has elapsed after selecting the one video by the reaction-based switching.
12. a word detection unit for detecting a specific word from the speech, 2. The switching device according to claim 1, wherein the switching control unit generates a switching signal for selecting a pre-prepared VTR video when the specific word is detected.
13. a tilt determination unit that receives a sensor signal indicating a tilt of a flip operated by at least one of the plurality of persons from a sensor provided on the flip and determines whether or not the tilt of the flip has exceeded a predetermined angle; The switching device according to claim 1 , wherein the switching control unit generates a switching signal for selecting a flip image obtained by capturing the flip when an inclination of the flip exceeds a predetermined angle.
14. A switching device according to claim 1; a switcher that receives the individual images of the plurality of persons and the switching signal and outputs the one image based on the switching signal; A switching system comprising:
15. Acquiring individual images of a plurality of persons and audio of the plurality of persons; generating a switching signal for selecting one image from a plurality of images including the individual images of the plurality of people based on at least one of the individual images and the audio; A program that causes a computer to execute a process of outputting the switching signal.
Citation Information
Patent Citations
Broadcast program production apparatus
JP2005229550A
Program generating system, command generating apparatus, and program generating program
JP2005295431A
Photographic shot controller, cooperative photographing system and program therefor
JP2010074327A
Automatic switching device, automatic switching method, and program
JP2022091640A
Device for controlling photographic shot and program for controlling photographic shot
JP2008072702A