Video generation system, video generation method, transmission device, and reception device
The video generation system addresses the loss of information in image transmission by generating and adjusting text data based on video conditions, ensuring accurate restoration of desired information in the receiving device.
Patent Information
- Application Number
- JP2024110026
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2026-01-22
AI Technical Summary
Existing image transmission methods that generate and transmit text data from video may result in the loss of required information during the restoration process.
A video generation system comprising a transmitting device that generates text data based on video content and predetermined conditions, and a receiving device that adjusts text generation conditions based on differences between the original and restored video, allowing for the transmission of video when specific conditions are met.
This system increases the likelihood that the restored video contains desired information by adjusting text generation conditions according to changes in purpose, content, or environment, thereby enhancing the accuracy of the video restoration process.
Smart Images

Figure 2026010284000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image generation system, an image generation method, a transmitting device, and a receiving device. [Background technology]
[0002] When transmitting images between multiple devices, it is desirable to reduce transmission capacity. There is a technology that reduces transmission capacity by having a transmitting device generate text data related to the video and transmit the text data to a receiving device, and then having the receiving device restore the video from the text data. For example, there is a technology that extracts image features from an image and generates text related to the image based on the image features (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2020 / 218111 Summary of the Invention [Problem to be solved by the invention]
[0004] However, when a transmitting device generates text data from video, transmits the text data to a receiving device, and the receiving device restores the video from the text data, the video generated may be missing information required by the receiving device.
[0005] An object of the present invention is to provide an image generation system, an image generation method, a transmitting device, and a receiving device that are capable of transmitting text data that increases the likelihood that an image containing desired information will be restored. [Means for solving the problem]
[0006] A video generation system based on the present disclosure includes a transmitting device and a receiving device that receives data from the transmitting device, wherein the transmitting device includes a text generation means that generates text data indicating the content of a first video based on a first video and text generation conditions, and a transmission means that transmits the generated text data to the receiving device, wherein the transmission means transmits the first video to the receiving device when a predetermined condition is met, and the receiving device includes a video generation means that generates a second video based on the text data received from the transmitting device, a condition generation means that generates text generation conditions based on a difference between the first video and the second video received from the transmitting device, and a condition notification means that notifies the transmitting device of the generated text generation conditions.
[0007] A video generation method based on the present disclosure includes a transmitting device generating text data indicating the content of a first video based on a first video and text generation conditions, the transmitting device transmitting the generated text data to a receiving device, the transmitting device transmitting the first video to the receiving device when a predetermined condition is met, the receiving device generating a second video based on the text data received from the transmitting device, the receiving device generating text generation conditions based on the difference between the first video and the second video received from the transmitting device, and the receiving device notifying the transmitting device of the generated text generation conditions.
[0008] A transmitting device based on the present disclosure includes a text generation means for generating text data indicating the content of the first video based on a first video and text generation conditions generated based on text data received by the receiving device, and a transmission means for transmitting the generated text data to the receiving device, and the transmission means transmits the first video to the receiving device when a predetermined condition is met.
[0009] A receiving device based on the present disclosure includes a video generation means for generating a second video based on text data indicating the content of a first video received from a transmitting device, a condition generation means for generating text generation conditions for the first video received from the transmitting device, and a condition notification means for notifying the transmitting device of the generated text generation conditions. [Effects of the Invention]
[0010] According to the present invention, it is possible to increase the possibility that a video containing desired information will be restored. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a block diagram illustrating an example of the configuration of a video generation system. [Figure 2] 10 is a flowchart illustrating an example of the operation of a transmission device in the video generation system. [Figure 3] 10 is a flowchart illustrating an example of the operation of a receiving device in the video generation system. [Figure 4] FIG. 1 is an explanatory diagram illustrating an example of the hardware configuration of an image generation system according to the present disclosure. [Figure 5] 10 is a flowchart illustrating an example of the operation of a transmission device in the video generation system. [Figure 6] 10 is a flowchart illustrating an example of the operation of a transmission device in the video generation system. [Figure 7] 10 is a flowchart illustrating an example of the operation of a transmission device in the video generation system. [Figure 8] 10 is a flowchart illustrating an example of the operation of a receiving device in the video generation system. [Figure 9] FIG. 1 is a block diagram illustrating an example of the configuration of a video generation system. [Figure 10] 10 is a flowchart illustrating an example of the operation of a transmission device in the video generation system. [Figure 11] FIG. 1 is a block diagram illustrating an overview of a video generation system. DETAILED DESCRIPTION OF THE INVENTION
[0012] Embodiment 1. In this embodiment, the video generation system includes a transmitting device and a receiving device that receives data from the transmitting device. The transmitting device generates text data indicating the content of the first video based on the first video and text generation conditions, transmits the generated text data to the receiving device, and transmits the first video to the receiving device when a predetermined condition is met. The receiving device generates a second video based on the text data received from the transmitting device, generates text generation conditions based on a difference between the first video and the second video received from the transmitting device, and notifies the transmitting device of the generated text generation conditions.
[0013] FIG. 1 is a block diagram showing an example of the configuration of a video generation system.
[0014] As shown in FIG. 1, the video generation system 100 includes a transmitting device 200, a receiving device 300 that receives data from the transmitting device 200, and a camera 400.
[0015] The transmitting device 200 includes a video acquisition unit 201, a text generation unit 202, a transmission unit 203, and a condition acquisition unit 204. The receiving device 300 includes a video / text acquisition unit 301, a video generation unit 302, a condition generation unit 303, a condition notification unit 304, and an analysis unit 305. The camera 400 is connected to the transmitting device 200. An image captured by the camera 400 is transmitted to the transmitting device 200.
[0016] The video acquisition unit 201 of the transmission device 200 acquires from the camera 400 the video captured by the camera 400 .
[0017] The text generation unit 202 of the transmitting device 200 creates (generates) text data indicating the content of the first video based on the video acquired by the video acquisition unit 201 (hereinafter referred to as the "first video") and the text generation conditions. In this example, the text generation conditions include both a first condition for including objects contained in the first video in the text data and a second condition for generating desired text data from objects detected from the first video.
[0018] For example, a first condition is set for detecting an object based on the number of pixels, such as "detect a group of 100 pixels or more in an image as an object." Specific examples of objects include people, vehicles, animals, buildings, etc. Furthermore, the text generation unit 202 can also convert visual information (position, color, etc.) of the detected object into text.
[0019] Furthermore, for example, a condition for generating text data from a detected object is set as a second condition, such as "generate text data using 30 Japanese characters or less."
[0020] The transmission unit 203 of the transmitting device 200 transmits the text data created by the text generation unit 202 to the receiving device 300. Specifically, the transmission unit 203 outputs the text data to a transmission path connecting the transmitting device 200 and the receiving device 300. Furthermore, the transmission unit 203 of the transmitting device 200 transmits the first video acquired by the video acquisition unit 201 to the receiving device 300 when a predetermined condition is met. Specifically, the transmission unit 202 outputs the first video to the transmission path connecting the transmitting unit 200 and the receiving unit 300 when a predetermined condition is met.
[0021] In this embodiment, the following six conditions are set as predetermined conditions for the transmitting device 200 to transmit the first video to the receiving device 300. First, there is little communication traffic in the transmission path for transmitting data from the transmitting device 200 to the receiving device 300. Second, there is a change in the purpose of using the second video. Third, there is a change of more than a predetermined amount in the content of the first video. For example, this applies when a predetermined number or a predetermined percentage of pixels in the video change. Fourth, there is a change in the shooting environment of the first video. For example, this applies when the shooting equipment such as the camera 400 is changed. Fifth, there is a low processing load on the transmitting device 200. Sixth, one or more of the following predetermined times that are set periodically occur. If at least one of these six conditions is met, the transmitting device 200 transmits the first video to the receiving device 300.
[0022] The video / text acquisition unit 301 of the receiving device 300 acquires the first video and text data transmitted from the transmission unit 202 of the transmitting device 200 .
[0023] The video generating unit 302 of the receiving device 300 generates a second video of the content indicated by the text data acquired by the video / text acquiring unit 301 .
[0024] The condition generation unit 303 of the receiving device 300 compares the first video acquired by the video / text acquisition unit 301 with the second video generated by the video generation unit 302, extracts differences between the first and second videos, and generates text generation conditions based on the differences.
[0025] The condition generating unit 303 may generate text generation conditions for making the second video closer to the user's desired image, on the premise that the user inputs necessary information into the receiving device 300. Alternatively, the condition generating unit 303 may generate text generation conditions so as to reduce the difference between the first video and the second video (in other words, so as to increase the similarity of the second video to the first video), on the premise that the user does not input necessary information into the receiving device 300.
[0026] For example, the condition generation unit 303 selects necessary elements from among elements such as the type of object (e.g., animal, human, vehicle, etc.), color, position, orientation, number, and attributes (e.g., gender, age group, breed of dog, etc.) as text generation conditions, and determines specific values of the selected elements.
[0027] The text generation conditions may be set to specify characteristics of the video. For example, the text generation conditions may be set so that the second video has specific characteristics, such as live action, animation, color, or black and white.
[0028] The condition notification unit 304 of the receiving device 300 notifies (transmits) the text generation conditions generated by the condition generation unit 303 to the transmitting device 200.
[0029] The condition acquisition unit 204 of the transmitting device 200 acquires the text generation conditions notified by the condition notification unit 304 of the receiving device 300. The text generation conditions acquired here are used when the sentence generation unit 202 generates text data, as described above.
[0030] Here, video and text data will be described with specific examples. First, as processing on the transmitting device 200 side, for example, the video acquisition unit 201 acquires, as a first video, a video including one pig and two trees as objects. Next, the sentence generation unit 202 generates text data of "one pig" based on the text generation condition. The text generation condition at this time is assumed to be "to convert the type and number of living things into text." Then, the transmission unit 203 transmits the text data and the first video.
[0031] Shifting to processing on the receiving device 300 side, the video / text acquisition unit 301 acquires the text data and the first video. Next, the video generation unit 302 generates a second video including one pig based on the text data. Then, the condition generation unit 303 extracts the number of trees as the difference between the first and second videos, and generates a new text generation condition as "convert the types and number of living things and the number of trees into text." The condition notification unit 304 notifies the transmitting device 200 of this new text generation condition.
[0032] The process returns to the transmitting device 200 side, where the condition acquisition unit 204 acquires new text generation conditions. The sentence generation unit 202 then uses the new text generation conditions the next time it generates text data. This allows adjustments to be made so that text data with the content desired by the receiving device 300 side is generated.
[0033] FIG. 2 is a flowchart showing an example of the operation of the transmission device 200 in the video production system 100.
[0034] In the transmitting device 200, the video acquisition unit 201 acquires the first video from the camera 400 (step S112).
[0035] In the transmitting device 200, the sentence generating unit 202 generates text data indicating the content of the first video based on the acquired first video and the text generation conditions (step S114). For example, the sentence generating unit 202 generates the text data from the first video using a generation AI (Artificial Intelligence) such as a VLM (Vision and Language Model).
[0036] The text generation conditions are stored in a storage area (not shown) of the transmitting device 200 each time they are acquired by the condition acquisition unit 204. When the sentence generation unit 202 generates text data for the first time, the text generation conditions stored in advance in a storage area (not shown) of the transmitting device 200 may be used. Alternatively, the transmitting device 200 may acquire the text generation conditions from the receiving device 300 when the video generation system 100 starts operating (when the transmitting device 200 and the receiving device 300 are connected), and generate the initial text data using the text generation conditions.
[0037] The transmitting device 200 determines whether or not a predetermined condition is satisfied (step S116). As described above, in this embodiment, the predetermined condition is one or more of the following: communication traffic in a transmission path for transmitting data from the transmitting device 200 to the receiving device 300 is less than a predetermined value; the purpose of using the second video has changed; the content of the first video has changed by a predetermined amount or more; the shooting environment of the first video has changed; the processing load of the transmitting device 200 is low; and a predetermined time that is set periodically has arrived.
[0038] If the predetermined condition is not satisfied, the transmission unit 203 in the transmission device 200 transmits the text data generated in step S114 to the reception device 300 (step S118). At this time, the transmission unit 203 does not transmit the primary video to the reception device 300.
[0039] If the predetermined condition is satisfied, transmission unit 203 in transmitting device 200 transmits the first video acquired in step S112 and the text data generated in step S114 to receiving device 300 (step S120). That is, if the predetermined condition is not satisfied, transmission unit 203 in transmitting device 200 transmits the text data to receiving device 300, but if the predetermined condition is satisfied, transmission unit 203 in transmitting device 200 transmits the first video in addition to the text data to receiving device 300.
[0040] Thereafter, the condition acquisition unit 204 in the transmitting device 200 acquires the text generation conditions notified (transmitted) from the condition notification unit 304 in the receiving device 300, and stores them in a storage area (not shown) of the transmitting device 200 (step S122). The text generation conditions stored here are used the next time the sentence generation unit 202 generates text data.
[0041] FIG. 3 is a flowchart showing an example of the operation of the receiving device 300 in the video production system 100.
[0042] In the receiving device 300, the video / text acquiring unit 301 acquires only the text data or both the text data and the primary video from the transmitting unit 203 of the transmitting device 200 (step S212).
[0043] In the receiving device 300, the video generation unit 302 generates a second video based on the acquired text data (step S214). For example, the video generation unit 302 generates the second video from the text data using a generation AI such as VLM.
[0044] The receiving device 300 determines whether or not the first video has been acquired in step S212 (step S216). If the first video has not been acquired, that is, if only text data has been acquired, the analysis unit 305 performs predetermined processing (for example, traffic congestion analysis processing, accident analysis processing, etc.) using the generated second video (step S218).
[0045] If it is determined that the first video has been acquired, that is, if both the text data and the first video have been acquired, the condition generation unit 303 extracts the differences between the acquired first video and the generated second video, and generates text generation conditions based on those differences (step S220).
[0046] In the receiving device 300, the condition notification unit 304 notifies the transmitting device 200 of the generated text generation conditions (step S222).
[0047] The following describes a specific example of the hardware configuration of the image generation system 100. Fig. 4 is an explanatory diagram showing an example of the hardware configuration of the image generation system according to the present disclosure.
[0048] The video production system 100 shown in FIG. 4 includes a camera 400, a transmitting device 200 connected to the camera 400, and a receiving device 300 connected to the transmitting device 200.
[0049] The transmission device 200 shown in FIG. 4 includes a CPU (Central Processing Unit) 1000, a main memory unit 1001, and an auxiliary memory unit 1002.
[0050] The transmitting device 200 is realized by software when the CPU 1000 shown in FIG. 4 executes a program that provides the functions of each component.
[0051] That is, the CPU 1000 loads a program stored in the auxiliary storage unit 1002 into the main storage unit 1001, executes it, and controls the operation of the transmission device 200, thereby realizing each function by software.
[0052] The main memory unit 1001 is used as a data working area and a data temporary saving area, and is, for example, a RAM (Random Access Memory).
[0053] The auxiliary storage unit 1002 is a non-transitory tangible storage medium, such as a magnetic disk, a magneto-optical disk, a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), or a semiconductor memory.
[0054] In the transmission device 200 of this embodiment, the auxiliary storage unit 1002 stores programs for realizing the video acquisition unit 201, the text generation unit 202, the transmission unit 203, and the condition acquisition unit 204.
[0055] The signal processing system may be implemented with a circuit including hardware components such as an LSI (Large Scale Integration) that realizes the functions shown in FIG.
[0056] The receiving device 300 shown in FIG. 4 includes a CPU (Central Processing Unit) 1020, a main memory unit 1021, and an auxiliary memory unit 1022.
[0057] Receiving device 300 is realized by software when CPU 1020 shown in FIG. 4 executes a program that provides the functions of each component.
[0058] That is, the CPU 1020 loads a program stored in the auxiliary storage unit 1022 into the main storage unit 1021, executes it, and controls the operation of the receiving device 300, thereby realizing each function by software.
[0059] The main memory unit 1021 is used as a data working area and a data temporary saving area, and is, for example, a RAM (Random Access Memory).
[0060] The auxiliary storage unit 1022 is a non-transitory tangible storage medium, such as a magnetic disk, a magneto-optical disk, a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), or a semiconductor memory.
[0061] In the receiving device 300 of this embodiment, the auxiliary storage unit 1022 stores programs for realizing the video / text acquisition unit 301, video generation unit 302, condition generation unit 303, condition notification unit 304, and analysis unit 305.
[0062] The signal processing system may be implemented with a circuit including hardware components such as an LSI (Large Scale Integration) that realizes the functions shown in FIG.
[0063] As described above, in this embodiment, the video generation system 100 includes the transmitting device 200 and the receiving device 300 that receives data from the transmitting device 200. The transmitting device 200 generates text data indicating the contents of the first video based on the first video and text generation conditions, transmits the generated text data to the receiving device 300, and transmits the first video to the receiving device when a predetermined condition is met. The receiving device 300 generates a second video based on the text data received from the transmitting device 200, generates text generation conditions based on the difference between the first video and the second video received from the transmitting device 200, and notifies the transmitting device 200 of the generated text generation conditions. This increases the likelihood that a video containing desired information will be restored.
[0064] The predetermined conditions are one or more of the following: low communication traffic on the transmission path for transmitting data from transmitting device 200 to receiving device 300; a change in the purpose of using the second video; a change in the content of the first video of a predetermined amount or more; a change in the shooting environment of the first video; a low processing load on transmitting device 200; and a predetermined, periodically set time. This allows the text generation conditions to be adjusted under appropriate circumstances.
[0065] For example, in a situation where there is a lot of communication traffic on the transmission path transmitting data from the transmitting device 200 to the receiving device 300, by transmitting the first video in addition to the text data from the transmitting device 200 to the receiving device 300, it is possible to generate text generation conditions without straining the transmission capacity.
[0066] For example, when the purpose of use of the second video has changed, by transmitting the first video in addition to the text data from the transmitting device 200 to the receiving device 300, it is possible to generate text generation conditions according to the changed purpose of use.
[0067] For example, in a situation where the content of the first video has changed by more than a predetermined amount, by transmitting the first video in addition to the text data from the transmitting device 200 to the receiving device 300, it is possible to generate text generation conditions that correspond to the change in the content of the first video.
[0068] For example, when a change occurs in the shooting environment of the first video, the transmitting device 200 transmits the first video in addition to the text data to the receiving device 300, thereby generating text generation conditions according to the changed shooting environment.
[0069] For example, when the processing load on the transmitting device 200 is low, by transmitting the first video in addition to the text data from the transmitting device 200 to the receiving device 300, it is possible to generate text generation conditions without increasing the processing load on the transmitting device 200.
[0070] For example, by transmitting the first video in addition to the text data from the transmitting device 200 to the receiving device 300 at a certain predetermined time that is set periodically, the text generation conditions can be adjusted periodically.
[0071] The text generation conditions include at least one or both of a condition for including a specific object related to the second video contained in the first video in the text data and a condition for generating desired text data from an object detected in the first video, thereby enabling the generation of text generation conditions with appropriate content.
[0072] Embodiment 2. In this embodiment, a transmission device generates a plurality of text data indicating the contents of each of a plurality of first videos based on a plurality of first videos and text generation conditions. The video generation system 100 of this embodiment is configured as shown in Fig. 1. However, the sentence generation unit 202 performs processing different from that of the first embodiment.
[0073] FIG. 5 is a flowchart showing an example of the operation of the transmission device 200 in the video production system 100.
[0074] In the transmitting device 200, the video acquisition unit 201 acquires a plurality of first videos from the camera 400 (step S132).
[0075] In the transmitting device 200, the text generating unit 202 generates a plurality of pieces of text data indicating the contents of each of the primary videos based on the acquired plurality of primary videos and the text generation conditions (step S134). Next, the transmitting device 200 performs the processes from step S116 onwards.
[0076] As described above, in this embodiment, the transmitting device 200 in the video production system 100 generates multiple pieces of text data indicating the contents of each of the multiple first videos based on the multiple first videos and the text generation conditions. This allows the receiving device 300 to generate multiple videos containing information required by the receiving device 300 from the text data received from the transmitting device 200.
[0077] Embodiment 3. In this embodiment, if the first video contains private data, the transmitting device generates text data excluding the private data. The video generation system 100 of this embodiment is configured as shown in Fig. 1. However, the transmitting device 200 performs processing slightly different from that of the first embodiment.
[0078] FIG. 6 is a flowchart showing an example of the operation of the transmission device 200 in the video production system 100.
[0079] After step S112, the transmitting device 200 determines whether or not privacy information is included in the acquired primary video (step S142).
[0080] If the transmitting device 200 determines that the acquired primary video does not include privacy information, it performs the processes in step S114 and thereafter.
[0081] When transmitting device 200 determines that the acquired first video contains privacy information, text generating unit 202 generates text data indicating the content of the first video excluding the privacy information based on the first video and the text generation conditions (step S144). Thereafter, transmitting device 200 performs the processes from step S116 onwards.
[0082] As described above, in this embodiment, when the first video includes privacy data, the transmitting device 200 in the video production system 100 generates text data excluding the privacy data, thereby enabling the receiving device 300 to generate the second video that does not include the privacy data.
[0083] Embodiment 4. In this embodiment, when the transmitting device is notified of the text generation conditions, the transmitting device regenerates text data based on the text generation conditions and transmits the regenerated text data to the receiving device, which then generates a second video based on the regenerated text data. The video generation system 100 of this embodiment is configured as shown in FIG. 1. However, the text generation unit 202 performs processing different from that of the first embodiment. Also, the video / text acquisition unit 301 performs processing slightly different from that of the first embodiment.
[0084] FIG. 7 is a flowchart showing an example of the operation of the transmission device 200 in the video production system 100.
[0085] After step S122, the text generation unit 202 in the transmission device 200 regenerates text data indicating the content of the primary video based on the acquired primary video and the new text generation conditions (step S152).
[0086] The transmission unit 203 in the transmission device 200 transmits the text data generated again in step S152 to the reception device 300 (step S154).
[0087] FIG. 8 is a flowchart showing an example of the operation of the receiving device 300 in the video production system 100.
[0088] After step S222, in the receiving device 300, the video / text acquiring unit 301 acquires the text data again from the transmitting unit 203 of the transmitting device 200 (step S252).
[0089] In the receiving device 300, the video generation unit 302 regenerates the second video based on the acquired text data (step S254). Then, the process proceeds to step S218. That is, the receiving device 300 executes a predetermined process using the regenerated second video.
[0090] As described above, in this embodiment, when the transmitting device 200 in the video production system 100 is notified of the text generation conditions, the transmitting device 200 regenerates text data based on the text generation conditions and transmits the regenerated text data to the receiving device 300. The receiving device 300 generates a second video based on the regenerated text data. This allows the receiving device 300 to generate a second video that reflects the updated text generation conditions.
[0091] Embodiment 5. In this embodiment, the video generation system includes multiple receiving devices, and the text generation means of the transmitting device generates common text data to be transmitted to the multiple receiving devices based on text generation conditions received from the multiple receiving devices.
[0092] FIG. 9 is a block diagram showing an example of the configuration of a video generation system.
[0093] The video generation system 110 of this embodiment includes a transmitting device 200, n (n≧2) receiving devices 300-1 to 300-n that receive data from the transmitting device 200, and a camera 400. The camera 400 is connected to the transmitting device 200. An image captured by the camera 400 is transmitted to the transmitting device 200. The transmitting device 200 is configured as shown in FIG.
[0094] Receiving devices 300-1 to 300-n have the same configuration as receiving device 300 shown in Fig. 1. The video generation unit in at least some of the multiple receiving devices 300-1 to 300-n generates video using an engine different from that used in the other receiving devices.
[0095] FIG. 10 is a flowchart showing an example of the operation of the transmission device 200 in the video production system 100.
[0096] In the transmitting device 200, after step S112, the sentence generating unit 202 generates common text data indicating the content of the first video based on the acquired first video and text generation conditions (step S162). The text generation conditions used here are n text generation conditions acquired from each of the multiple receiving devices 300-1 to 300-n. The common text data is text data adjusted so that a common second video is generated in the multiple receiving devices 300-1 to 300-n. Specifically, the transmitting device 200 acquires engine information (information indicating what engine is operating the video generation unit, for example, engine type information) from each of the receiving devices 300-1 to 300-n. Then, the sentence generating unit 202 analyzes the engine information and generates, as common text data, text data adjusted so that a common second video is generated in all of the engines. The common text data is generated by the transmitting device 200 and transmitted to the plurality of receiving devices 300-1 to 300-n, so that a common second video is generated even if the engines of the video generating units in the plurality of receiving devices are different.
[0097] After step S162, the transmission device 200 proceeds to step S116. If the predetermined condition is not satisfied in step S116, the transmission unit 203 in the transmission device 200 transmits the common text data generated in step S162 to each of the reception devices 300-1 to 300-n (step S164).
[0098] If the predetermined condition is met in step S116, transmission unit 203 in transmission device 200 transmits the primary video acquired in step S112 and the common text data generated in step S162 to each of reception devices 300-1 to 300-n (step S166).
[0099] Thereafter, the condition acquisition unit 204 in the transmitting device 200 acquires the text generation conditions notified (transmitted) from the condition notification unit 304 in each of the receiving devices 300-1 to 300-n, and stores them in a storage area (not shown) of the transmitting device 200 (step S168). The text generation conditions stored here are used the next time the sentence generation unit 202 generates text data.
[0100] As described above, in this embodiment, the video production system 100 includes a plurality of receiving devices 300-1 to 300-n, and the text generation unit 202 of the transmitting device 200 generates common text data to be transmitted to the plurality of receiving devices 300-1 to 300-n based on the text generation conditions received from the plurality of receiving devices 300-1 to 300-n. This allows a common second video to be generated in the plurality of receiving devices.
[0101] In this embodiment, the transmitting device 200 performs a process of adjusting the plurality of receiving devices 300-1 to 300-n so that a common second video is generated in the plurality of receiving devices 300-1 to 300-n, but the present invention is not limited to this. For example, a specific receiving device among the plurality of receiving devices 300-1 to 300-n may acquire engine information of the video generating units in the plurality of receiving devices 300-1 to 300-n. The specific receiving device may then analyze the engine information to generate common text generation conditions adjusted so that a common second video is generated in all of the engines, and notify the transmitting device 200 of the common text generation conditions. The transmitting device 200 may generate common text data by generating text data based on the common text generation conditions.
[0102] Next, an overview of the present invention will be described. FIG. 11 is a block diagram showing an overview of a video generation system according to the present invention. The video generation system 10 according to the present invention includes a transmitting device 20 and a receiving device 30 that receives data from the transmitting device. The transmitting device 20 includes a text generation means 21 that generates text data indicating the content of a first video based on a first video and text generation conditions, and a transmission means 22 that transmits the generated text data to the receiving device. The transmission means 22 transmits the first video to the receiving device 30 when a predetermined condition is met. The receiving device 30 includes a video generation means 31 that generates a second video based on the text data received from the transmitting device 20, and a condition generation means 32 that generates a text generation condition based on a difference between the first video and the second video received from the transmitting device 20. The receiving device 30 also includes a condition notification means 33 that notifies the transmitting device 20 of the generated text generation condition. This configuration increases the likelihood of restoring a video containing desired information.
[0103] The predetermined condition may be one or more of the following: low communication traffic on the transmission path for transmitting data from transmitting device 20 to receiving device 30; a change in the purpose of using the second video; a change in the content of the first video by a predetermined amount or more; a change in the shooting environment of the first video; a low processing load on transmitting device 20; and a predetermined, periodically set time. Such a configuration allows the text generation conditions to be adjusted under appropriate circumstances.
[0104] The text generation conditions may include at least one or both of a condition for including a specific object related to the second video contained in the first video in the text data and a condition for generating desired text data from an object detected in the first video, thereby enabling the generation of text generation conditions with appropriate content.
[0105] Furthermore, the text generation means 21 may generate a plurality of text data indicating the contents of each of the plurality of first videos based on the plurality of first videos and the text generation conditions. This allows a plurality of videos including information required by the receiving device 30 to be generated from the text data received from the transmitting device 20.
[0106] Furthermore, if the first video contains private data, the text generation means 21 may generate text data excluding the private data, thereby enabling the receiving device 30 to generate the second video that does not contain the private data.
[0107] Furthermore, when notified of the text generation conditions, the sentence generation means 21 may regenerate text data based on the text generation conditions. The transmission means 22 may transmit the regenerated text data to the receiving device, and the image generation means 31 may generate a second image based on the regenerated text data. This allows the receiving device 30 to generate a second image that reflects the updated text generation conditions.
[0108] Furthermore, a plurality of receiving devices 30 may be provided, and the text generation means 21 of the transmitting device 20 may generate common text data to be transmitted to the plurality of receiving devices 30 in common, based on the text generation conditions received from the plurality of receiving devices 30. This allows a common second video to be generated in the plurality of receiving devices.
[0109] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.
[0110] (Appendix 1) A transmitting device; a receiving device that receives data from the transmitting device, The transmitting device a text generation means for generating text data indicating the content of the first video based on the first video and a text generation condition; a transmission means for transmitting the generated text data to the receiving device, the transmitting means transmits the first video to the receiving device when a predetermined condition is met; The receiving device an image generating means for generating a second image based on the text data received from the transmitting device; a condition generating means for generating the text generation condition based on a difference between the first image and the second image received from the transmitting device; and a condition notification means for notifying the transmission device of the generated text generation conditions. Image generation system.
[0111] (Appendix 2) The predetermined conditions are: The occurrence of one or more of the following events occurs: there is little communication traffic in a transmission path for transmitting data from the transmitting device to the receiving device; the purpose of using the second video has changed; there has been a change in the content of the first video by a predetermined amount or more; there has been a change in the shooting environment of the first video; the processing load of the transmitting device is low; and a certain predetermined time that is set periodically has arrived. 10. The video generation system of claim 1.
[0112] (Supplementary Note 3) The text generation conditions include at least one or both of a condition for including a specific object related to the second video included in the first video in the text data and a condition for generating desired text data from an object detected in the first video. 10. The video generation system of claim 1.
[0113] (Supplementary Note 4) The sentence generation means generates a plurality of text data indicating the contents of each of the plurality of first videos based on the plurality of first videos and text generation conditions. 4. The video generation system according to any one of claims 1 to 3.
[0114] (Supplementary Note 5) When the first video includes private data, the sentence generation means generates the text data excluding the private data. 4. The video generation system according to any one of claims 1 to 3.
[0115] (Supplementary Note 6) When the text generation conditions are notified, the sentence generation means regenerates text data based on the text generation conditions; the transmitting means transmits the regenerated text data to the receiving device; The image generating means generates a second image based on the regenerated text data. 4. The video generation system according to any one of claims 1 to 3.
[0116] (Supplementary Note 7) A plurality of the receiving devices are provided, the text generation means of the transmitting device generates common text data to be commonly transmitted to the plurality of receiving devices based on the text generation conditions received from the plurality of receiving devices; 4. The video generation system according to any one of claims 1 to 3.
[0117] (Supplementary Note 8) A transmitting device generates text data indicating the content of the first video based on the first video and a text generation condition; the transmitting device transmits the generated text data to the receiving device; the transmitting device transmits the first video to the receiving device when a predetermined condition is met; the receiving device generates a second image based on the text data received from the transmitting device; the receiving device generates the text generation condition based on a difference between the first image and the second image received from the transmitting device; The receiving device notifies the transmitting device of the generated text generation conditions. Video generation method.
[0118] (Supplementary Note 9) A sentence generation means for generating the text data indicating the content of the first video based on a first video and a text generation condition generated based on text data received by a receiving device; a transmission means for transmitting the generated text data to the receiving device, The transmitting means transmits the first video to the receiving device when a predetermined condition is met. Transmitting device.
[0119] (Supplementary Note 10) A video generation means for generating a second video based on text data indicating the content of the first video received from the transmitting device; a condition generating means for generating a text generation condition for the first video received from the transmitting device; and a condition notification means for notifying the transmission device of the generated text generation conditions. Receiving device.
[0120] Furthermore, some or all of the configurations described in Supplementary Notes 2 to 7, which are dependent on Supplementary Note 1, may also be dependent on Supplementary Notes 8 to 10 in the same dependent relationship as Supplementary Notes 2 to 7. Furthermore, not limited to Supplementary Notes 1 and 8 to 10, some or all of the configurations described as Supplements may be made dependent on various hardware, software, various recording means for recording software, or systems, within the scope of each of the above-mentioned embodiments. [Explanation of symbols]
[0121] 10,100,110 Image Creation System 20 Transmitting device 21 Sentence generation means 22 Transmission means 30 Receiving device 31 Image generation means 32 Condition generation means 33 Condition notification means 200 Transmitting device 201 Video Acquisition Unit 202 Sentence generation section 203 Transmission Unit 204 Condition Acquisition Section 300 Receiving device 300-1 to 300-n receiving device 301 Video and Text Acquisition Department 302 Image Generation Unit 303 Condition generator 304 Condition Notification Section 305 Analysis Department 400 cameras 1000 CPU 1001 Main memory 1002 Auxiliary storage 1020 CPU 1021 Main memory 1022 Auxiliary storage
Claims
1. a transmitting device; a receiving device that receives data from the transmitting device, The transmitting device a text generation means for generating text data indicating the content of the first video based on the first video and a text generation condition; a transmission means for transmitting the generated text data to the receiving device, the transmitting means transmits the first video to the receiving device when a predetermined condition is met; The receiving device an image generating means for generating a second image based on the text data received from the transmitting device; a condition generating means for generating the text generation condition based on a difference between the first image and the second image received from the transmitting device; and a condition notification means for notifying the transmission device of the generated text generation conditions. Image generation system.
2. The predetermined condition is: The occurrence of one or more of the following events occurs: there is little communication traffic in a transmission path for transmitting data from the transmitting device to the receiving device; the purpose of using the second video has changed; there has been a change in the content of the first video by a predetermined amount or more; there has been a change in the shooting environment of the first video; the processing load of the transmitting device is low; and a certain predetermined time that is set periodically has arrived. The video production system of claim 1 .
3. The text generation conditions include at least one or both of a condition for including a specific object related to the second video, which is included in the first video, in the text data and a condition for generating desired text data from an object detected from the first video. The video production system of claim 1 .
4. The sentence generation means generates a plurality of text data indicating the contents of each of the plurality of first videos based on a plurality of first videos and a text generation condition. The image generation system according to any one of claims 1 to 3.
5. When the first video includes private data, the text generation means generates the text data excluding the private data. The image generation system according to any one of claims 1 to 3.
6. the sentence generation means, when notified of the text generation conditions, regenerates text data based on the text generation conditions; the transmitting means transmits the regenerated text data to the receiving device; The image generating means generates a second image based on the regenerated text data. The image generation system according to any one of claims 1 to 3.
7. a plurality of the receiving devices; the text generation means of the transmitting device generates common text data to be commonly transmitted to the plurality of receiving devices based on the text generation conditions received from the plurality of receiving devices; The image generation system according to any one of claims 1 to 3.
8. a transmitting device generating text data indicating the content of the first video based on the first video and a text generation condition; the transmitting device transmits the generated text data to the receiving device; the transmitting device transmits the first video to the receiving device when a predetermined condition is met; the receiving device generates a second image based on the text data received from the transmitting device; the receiving device generates the text generation condition based on a difference between the first image and the second image received from the transmitting device; The receiving device notifies the transmitting device of the generated text generation conditions. Video generation method.
9. a text generation means for generating text data indicating the content of the first video based on a first video and a text generation condition generated based on text data received by a receiving device; a transmission means for transmitting the generated text data to the receiving device, The transmitting means transmits the first video to the receiving device when a predetermined condition is met. Transmitting device.
10. an image generating means for generating a second image based on text data indicating the content of the first image received from the transmitting device; a condition generating means for generating a text generation condition for the first video received from the transmitting device; and a condition notification means for notifying the transmission device of the generated text generation conditions. Receiving device.
Citation Information
Patent Citations
Learning method and device, program, learned model, and text generation device
WO2020218111A1