Image synthesis and distribution system and method
The video compositing and distribution system addresses delays in existing systems by compositing and transmitting partial images with identifiers, reducing bandwidth and processing time through selective packet handling.
Patent Information
- Application Number
- PCT/JP2024/013352
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-10-02
AI Technical Summary
Existing video distribution systems experience significant delays due to data compression and bandwidth pressure, particularly in technologies that require transmitting and decoding the base layer every time, leading to increased communication load and processing time.
A video compositing and distribution system that selects multiple partial images with different coordinates from a single image, composites them, and transmits these images with a common identifier, allowing selective reception and decoding to reduce communication bandwidth and delays.
This approach reduces communication bandwidth and delays by enabling selective packet transmission and reception, minimizing the need for decoding the entire base layer, and allows for flexible video quality adaptation with reduced processing time.
Smart Images

Figure JP2024013352_02102025_PF_FP_ABST
Abstract
Description
Video composition distribution system and method
[0001] The present disclosure relates to a video synthesis and distribution system that synthesizes a plurality of video signals.
[0002] In recent video distribution systems, distribution is performed using web technology. For example, in Non-Patent Document 1, compressed video is divided into small data series called segments, and compressed data for each video quality is prepared in advance, so that the receiving side can selectively receive video of a quality that corresponds to the capacity of the network bandwidth, etc.
[0003] Also known is a video distribution method using scalable coding in which a single video is hierarchically coded (see Non-Patent Documents 2 and 3). For example, Non-Patent Document 2 discloses a technology in which a video is coded by dividing it into a base layer and multiple enhancement layers, and low-quality video is reproduced when only the base layer is received, but the more enhancement layers are received in addition to the base layer, the higher-quality video is reproduced.
[0004] STOCKHAMMER, Thomas. Dynamic adaptive streaming over HTTP - standards and design principles. In: Proceedings of the second annual ACM conference on multimedia systems. 2011. pp. 133-144. Kumi Shinsenji, Kazuto Uekura, Shigetaro Iwatsu, Katsuhiko Fukasawa, Yoshiyuki Yashima: Scalable video distribution technology, NTT Technical Journal, July 2005, pp. 51-54. Internet <https: / / journal.ntt.co.jp / . jp / backnumber2 / 0507 / files / jn200507051.pdf> Takamura Masayuki, Bando Yukihiro, Shimizu Shinya; Chapter 10: High-Performance Coding Processing, "Knowledge Base" of the Institute of Electronics, Information and Communication Engineers, 2013 Internet <https: / / www.ieice-hbkb.org / files / 02 / 02gun_05hen_10.pdf>
[0005] However, the technologies disclosed in Non-Patent Documents 1 to 3 are intended for data exchange between frames involving data compression, and have a problem in that delays of several hundred milliseconds or more occur due to the codec used when synthesizing images. In particular, the technology disclosed in Non-Patent Document 2 requires the transmitting side to transmit the base layer every time, which puts pressure on the communication bandwidth and increases the load, causing delays. Furthermore, the receiving side requires the receiving side to receive and decode all of the base layer data every time, which takes a long time to process and increases the delay.
[0006] In order to solve the above-mentioned problems, an object of the present disclosure is to provide a video compositing and distribution system that can composite videos while suppressing delays.
[0007] In order to achieve the above object, the compositing and distribution device and video compositing and distribution system of the present disclosure employ a technique of selecting multiple partial images with different coordinates from a single image and compositing the multiple partial images.
[0008] Specifically, the video compositing and distribution system of the present disclosure comprises a sending-side compositing and distribution device that selects multiple partial images with different coordinates from a single image, composites the multiple partial images, and transmits them; and a receiving-side compositing and distribution device that decodes the image using data of the composited multiple partial images transmitted from the sending-side compositing and distribution device.
[0009] With this configuration, packets can be transmitted separately for each coordinate in the image, preventing duplication of the base layer in the packets and reducing the amount of transmitted data. Furthermore, uncompressed, scalably encoded, packetized video can be selectively received. This reduces the communication bandwidth and delays. This reduces the delay from video signal input to video output, in particular.
[0010] In addition, the sending-side synthesis and distribution device may assign a common identifier to the multiple partial images to be selected, and select multiple different partial images from the single image based on the identifier, and the receiving-side synthesis and distribution device may receive data of the multiple partial images based on the identifier and decode the images.
[0011] With this configuration, packets are selectively sent and received, which reduces communication bandwidth and delays. Furthermore, since it is sufficient to identify only a portion of the encoded identifier, the number of matching tables can be reduced. By appropriately setting the GID, reduced video can be acquired with low delay and a high degree of flexibility.
[0012] The identifier may rotate on the single image at a predetermined cycle.
[0013] With this configuration, packets are selectively sent and received, which reduces communication bandwidth and delays. Also, since it is sufficient to identify only a portion of the encoded identifier, the number of matching tables can be reduced.
[0014] Furthermore, the image may be a frame image included in the video, and the sending-side synthesis and distribution device may assign a common identifier to a frame image that rotates at a predetermined time period among a plurality of frame images included in the video at different times, and the receiving-side synthesis and distribution device may receive the frame image based on the identifier.
[0015] With this configuration, packets are selectively sent and received, which reduces communication bandwidth and delays. Also, since it is sufficient to identify only a portion of the encoded identifier, the number of matching tables can be reduced.
[0016] The identifier may have a bit length of two or more, the sending combining and distribution device may divide the identifier into two or more parts and store each divided identifier in a predetermined packet area, and the receiving combining and distribution device may obtain the identifier by reading the predetermined packet area.
[0017] This allows for a significant reduction in the number of filter rules.
[0018] In addition, the sending-side combining and distributing device may generate a packet for each of the same groups, encode the identifier, store it in a predetermined area of the packet, and selectively transmit the packet in response to a transmission request based on the identifier from the receiving-side combining and distributing device, and the receiving-side combining and distributing device may selectively receive the packet by identifying a part of the encoded identifier stored in the predetermined area, and decode the image using one or more of the packets.
[0019] This allows packets to be selectively sent and received, which reduces communication bandwidth and delays. Also, since it is sufficient to identify only a portion of the encoded identifier, the number of matching tables can be reduced.
[0020] The system may also include an interconnect interposed between the sending-side combining and distribution device and the receiving-side combining and distribution device, which receives the packet having the encoded identifier stored in the header portion from the receiving-side combining and distribution device in response to a transmission request based on the identifier of the receiving-side combining and distribution device, and selectively forwards the packet to the sending-side combining and distribution device.
[0021] This allows the use of general L2 and L3 switches as interconnects.
[0022] The present disclosure also provides a video synthesis distribution method, including the steps of: selecting a plurality of partial images having different coordinates from one image, synthesizing the plurality of partial images, and transmitting the selected partial images; and decoding the image using data of the synthesized plurality of partial images.
[0023] The device of the present invention can also be realized by a computer and a program, and the program can be recorded on a recording medium or provided via a network. The program of the present disclosure is a program for causing a computer to realize each function of the device according to the present disclosure, and a program for causing a computer to execute each procedure of the method executed by the device according to the present disclosure.
[0024] The above disclosures can be combined as much as possible.
[0025] According to the present disclosure, it is possible to synthesize videos while suppressing delays associated with data compression.
[0026] 1 is a diagram illustrating an overview of a video compositing and distribution system according to an embodiment of the present disclosure; FIG. 2 is a diagram illustrating a configuration of a compositing and distribution device according to an embodiment of the present disclosure; FIG. 3 is a diagram illustrating scalable coding; FIG. 4 is a diagram illustrating identification of a group ID using identifier coding; FIG. 5 is a diagram illustrating scale interleaving; FIG. 6 is a diagram illustrating a comparison between coding the GIDs of X and Y as they are and coding them with scale codes; FIG. 7 is a diagram illustrating a comparison between a filter rule when coding the GIDs of X and Y as they are and coding them with scale codes; FIG. 8 is a diagram illustrating the effect of scale interleaving; FIG. 9 is a diagram illustrating the configuration of a video packet when a GID is stored in an RTP-extended frame; FIG. 10 is a diagram illustrating the configuration of a video packet when a GID is stored in a destination MAC address; FIG. 11 is a diagram illustrating the configuration of a video packet when a GID is stored in a VLAN; and FIG. 12 is a diagram illustrating the configuration of a video packet when a GID is stored in a destination IP address.
[0027] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the present disclosure is not limited to the embodiments shown below. These implementation examples are merely illustrative, and the present disclosure can be implemented in various forms with various modifications and improvements based on the knowledge of those skilled in the art. Note that components with the same reference numerals in this specification and drawings indicate the same components.
[0028] [Overall Overview of Video Compositing and Distribution System] A video compositing and distribution system 300 according to an embodiment of the present disclosure will be described with reference to FIGS. 1 to 3. First, an overall overview of the video compositing and distribution system 300 will be described based on FIG. 1. The video compositing and distribution system 300 includes a control device 10, multiple combining and distribution devices 30A, 30B, 30C, and 30D, and an interconnect 50. In the following description, when no distinction is made between the combining and distribution devices 30A, 30B, 30C, and 30D, they will simply be referred to as combining and distribution devices 30. Furthermore, each of the multiple combining and distribution devices 30 functions as a "sending-side combining and distribution device" and also functions as a "receiving-side combining and distribution device."
[0029] The video compositing and distribution system 300 of the present disclosure employs scalable coding that circulates at a predetermined cycle. Specifically, a time axis (t) of the video and coordinates (x, y) are set for each time frame, and each of t, x, and y that circulate at a predetermined cycle is referred to as a group ID. Any combination of t, x, and y is also referred to as a group ID. The sending-side compositing and distribution device 30 then associates the group ID with an identifier consisting of a bit string and packetizes the video according to the combination of t, x, and y. The receiving-side compositing and distribution device 30 decodes the video by identifying a portion of the bit string. Note that in this embodiment, an image is divided into two coordinates, X and Y, but the scope of the present disclosure is not limited to this and includes any case in which an image is divided into one or more coordinates. The group ID is an example of an "identifier."
[0030] In the following embodiment, the coordinates for low-quality video are (t, x, y) = (0, 0, 0), and an example is shown in which high-quality video is constructed by interpolating the coordinates (0, 0, 0) with three coordinates (see FIG. 3). The coordinates (0, 0, 0) are arranged at intervals according to the resolution for the low-quality video, so one time frame contains the number of coordinates (0, 0, 0) according to the resolution for the low-quality video. The coordinates for high-quality video are also interpolated with three coordinates, and are repeated for each coordinate (0, 0, 0).
[0031] Specifically, the video compositing and distribution system 300 of the present disclosure includes a sending-side compositing and distribution device 30 that selects multiple partial images with different coordinates from one image, combines the multiple partial images, and transmits them; and a receiving-side compositing and distribution device 30 that decodes an image using data of the multiple composited partial images transmitted from the sending-side compositing and distribution device 30.
[0032] In addition, the sending-side composite distribution device 30 assigns a common group ID to multiple partial images to be selected, and selects multiple different partial images from one image based on the group ID, and the receiving-side composite distribution device 30 receives data on the multiple partial images based on the group ID and decodes the images.
[0033] Specifically, in the video compositing and distribution system 300 of the present disclosure, the transmitting-side compositing and distribution device 30 generates a packet for each identical group ID, encodes the group ID, stores it in a predetermined area of the packet, and selectively transmits the packet in response to a transmission request based on the group ID from the receiving-side compositing and distribution device, and the receiving-side compositing and distribution device selectively receives the packet by identifying a part of the encoded group ID stored in the predetermined area, and decodes an image using one or more packets.
[0034] Specifically, the video compositing and distribution system 300 of the present disclosure includes an interconnect 50 that is interposed between the sending-side compositing and distribution device 30 and the receiving-side compositing and distribution device 30, and that receives packets having an encoded group ID stored in the header portion from the receiving-side compositing and distribution device 30 in response to a transmission request based on the group ID of the receiving-side compositing and distribution device 30, and selectively transfers the packets to the sending-side compositing and distribution device 30.
[0035] The multiple combining and distributing devices 30 are connected to one another via an interconnect 50. Video can be input to the combining and distributing device 30 from an external source, and video can also be input via the interconnect 50. Video can also be input to the combining and distributing device 30 from a video distribution device external to the system. The combining and distributing device 30 transmits the input video. Furthermore, the combining and distributing device 30 generates video by combining video acquired from another combining and distributing device 30 or a video distribution device external to the system via the interconnect 50 with the video input to itself.
[0036] The combining and distributing device 30 packetizes and transmits uncompressed or lightly compressed video. At this time, the combining and distributing device 30 performs scalable encoding and packetization using a method described below. This makes it possible to selectively receive video with the required video resolution, image quality, and frame rate. Here, "light compression" refers to a compression method that is closed within a frame (still image) and does not perform inter-frame (temporal) compression such as MPEG (Moving Picture Experts Group).
[0037] When receiving video from another combining and distributing device 30, the combining and distributing device 30 can request the interconnect 50 via the control device 10 to transmit the necessary video according to its own capabilities. Here, the necessary video according to the capabilities of the receiving combining and distributing device 30 refers not only to the video data itself that is the target content, but also to video that has the necessary video resolution, image quality, and frame rate for the receiving combining and distributing device 30. This makes it possible to reduce the transmission bandwidth between the combining and distributing device 30 and the interconnect 50. It also makes it possible to minimize the scale of the receiving system of the combining and distributing device 30.
[0038] The scope of the present disclosure is not limited to the receiving combining / distributing device 30 requesting the interconnect 50 for the video resolution or image quality, frame rate, etc. required for itself, but the sending combining / distributing device 30 may also request the interconnect 50 for the video resolution or image quality, frame rate, etc. required for the receiving combining / distributing device 30.
[0039] The interconnect 50 is interposed between multiple combining and distribution devices 30, receives video from each combining and distribution device 30, and forwards the video to other combining and distribution devices 30. At this time, the interconnect 50 can selectively transmit video packets based on a request from a receiving or sending combining and distribution device 30 via the control device 10. For example, in response to a request from a combining and distribution device 30A based on a group ID described below, the interconnect 50 instructs one or more of the combining and distribution devices 30B, 30C, and 30D to selectively transmit video packets based on the group ID, and forwards the video packets received from the one or more combining and distribution devices 30 to the requesting combining and distribution device 30A.
[0040] The interconnect 50 may also sequentially accumulate video packets and transmit them in response to a transmission request from the combining and distribution device 30 based on the group ID. The interconnect 50 may be configured with multiple receiving transfer devices and a network connecting them. For example, it is assumed that a general L2 or L3 switch is used as the receiving transfer device. The network that constitutes the interconnect 50 may be wide-area. The receiving transfer device may also be configured to enable priority control of video packets.
[0041] In the video compositing and distribution system 300, the control device 10 is not necessarily an essential component, and the compositing and distribution device 30 may directly request the interconnect 50 to transmit the required video.
[0042] [Configuration of the Combining and Distribution Device] Next, the specific configuration of the combining and distribution device 30 will be described with reference to Figure 2. Each of the combining and distribution devices 30A, 30B, 30C, and 30D has the same configuration. The combining and distribution device 30 has a video input IF (Interface) 31, an encoding unit 32, a video combining unit 33, a video output IF 34, a network IF 35, a decoding unit 36, a video request unit 37, a network control unit 38, and a control receiving unit 39.
[0043] The video input IF 31 converts an externally input video signal into a video signal for internal processing, and passes the data to the encoding unit 32 and the video synthesis unit 33. A plurality of video signals may be input to the video input IF 31. For example, video signals from a plurality of locations may be input to the video input IF 31 substantially simultaneously.
[0044] The encoding unit 32 performs scalable encoding and packetization, as described below, on the video signal from the video input IF 31. The encoding unit 32 may perform scalable encoding and packetization on a plurality of video signals. For example, the encoding unit 32 may perform scalable encoding and packetization on video signals relating to video from a plurality of locations. The encoding unit 32 passes the packetized data to the network IF 35.
[0045] The network IF 35 transfers video packets from the encoding unit 32 to the outside, and transfers video packets from the outside to the decoding unit 36. The network IF 35 also mediates requests from a video request unit 37 and control information from a network control unit 38 to the outside. Here, mediation includes input of requests and information from the outside.
[0046] The decoding unit 36 receives the scalably encoded and packetized video packets and decodes them into a video signal. The decoding unit 36 may decode multiple video signals. For example, the decoding unit 36 may decode video signals related to video from multiple locations.
[0047] The video synthesis unit 33 synthesizes a new video from the video signal from the video input IF 31 and the video signal from the decoding unit 36, and outputs the synthesized video to the video output IF 34. The video synthesis method at this time follows instructions from the control receiving unit 39. Furthermore, if there is any video that is required other than the video input from the video input IF 31, the video synthesis unit 33 makes a request to the video request unit 37.
[0048] The video request unit 37 requests the necessary video packets via the network IF 35 based on the request from the video synthesis unit 33 and information from the network control unit 38. Specifically, the video request unit 37 requests the necessary video packets based on the following group ID. The control acceptance unit 39 accepts information on the screen layout of the synthesized video and the video source from the outside, and issues instructions to the video synthesis unit 33.
[0049] The video output IF 34 converts the internal video signal from the video synthesis unit 33 into a video signal for external processing and outputs the converted video signal. The video output IF 34 may output a plurality of video signals. For example, the video output IF 34 may output video signals from a plurality of locations.
[0050] The network control unit 38 manages management information required for processing in the encoding unit 32 and decoding unit 36 for each video, and provides the management information to the encoding unit 32, decoding unit 36, and video request unit 37. The network control unit 38 can also transmit the management information to the outside via the network IF 35. The network control unit 38 can also obtain, via the network IF 35, management information for videos processed in other combining and distribution devices 30.
[0051] [Scalable Code] Next, the scalable code used in the scalable coding performed by the coding unit 32 will be described with reference to FIG.
[0052] Specifically, the composite distribution device 30 of the present disclosure includes an encoding unit 32 that selects multiple partial images that are different from one another from a single image and composites the multiple partial images, and a decoding unit 36 that decodes the image using data of the multiple composite partial images.
[0053] In addition, the encoding unit 32 may assign a group ID that rotates on a single image to multiple partial images to be selected, and select multiple partial images that are different from each other from the single image based on the group ID.
[0054] The encoding unit 32 may also assign the group ID common to some images to select some images from a plurality of images linked to a time, select some images from a plurality of images linked to a time based on the group ID, and select multiple partial images with different coordinates for each selected image, and synthesize the partial images for the partial images, and the decoding unit 36 may decode the multiple images using data of the multiple partial images in each synthesized image.
[0055] Furthermore, the images may be frame images included in the video, and the encoding unit 32 may assign a common group ID to frame images that rotate at a predetermined time period among multiple frame images included in the video at different times, and the decoding unit 36 may select the frame images based on the group ID.
[0056] In addition, the group ID has a bit length of 2 or more, the encoding unit 32 divides the group ID into 2 or more parts and stores each divided group ID in a predetermined packet area, and the decoding unit 36 obtains the group ID by reading the predetermined packet area.
[0057] First, in this embodiment, a number ranging from t=0 to a maximum of T is assigned to each frame in a cyclic manner as a time-related group ID (GID (Group Identification)) on the time axis of the video. FIG. 3 shows the case where T=3. Furthermore, for each time frame, x and y coordinates are set, and as the coordinate-related group ID, a number ranging from 0 to a maximum X is assigned to x and a number ranging from 0 to a maximum Y is assigned to y in a cyclic manner. FIG. 3 shows the case where X=3 and Y=3. That is, in this embodiment, 4×4×4=64 levels of scaling are realized for the video. Note that, although the cyclic sequence length of each of t, x, and y in this embodiment is 4, it may be set to several tens or several hundreds. Here, the group ID does not necessarily need to be cyclic; it is sufficient that the encoding unit 32 is configured to be able to select a portion of the image. Note that when the group ID is cyclically assigned, it is not necessary to change the group ID for each pixel; a group unit can also be made up of multiple different pixels.
[0058] In other words, the video compositing and distribution system 300 assigns a group ID (t=0 to 3) to each time frame constituting the video, and manages those assigned the same group ID as a single group. The video compositing and distribution system 300 also sets an x-coordinate for each time frame, assigns a group ID (t=0 to 3) to each time frame, and manages those assigned the same group ID (a vertical column of the time frames in the figure) as a single group. The video compositing and distribution system 300 also sets a y-coordinate for each time frame, assigns a group ID (y=0 to 3) to each time frame, and manages those assigned the same group ID (a horizontal row of the time frames in the figure) as a single group. This cyclical assignment of group IDs up to T, X, and Y to the video time axis t and screen coordinates x and y is called scalable coding. The video compositing and distribution system 300 also manages any combination of t, x, and y by assigning the same group ID to the same combination.
[0059] The scope of the present disclosure is not limited to assigning a group ID to two-dimensional screen coordinates x and y as in the present embodiment, but may be configured to assign a group ID to each of three-dimensional coordinates x, y, and z.
[0060] [Packetization] The encoding unit 32 of the sending-side combining and distributing device 30 generates video packets for each group ID (Distributed Video Frame, packetization). The network control unit 38 manages video management information necessary for processing by the encoding unit 32 and decoding unit 36, such as the group ID and address information corresponding to each video, and provides the management information to the encoding unit 32, decoding unit 36, and video request unit 37. Note that the scope of the present disclosure is not limited to the encoding unit 32, which is a single component, performing both scalable encoding and packetization processing, and separate components may be configured to perform scalable encoding and packetization, respectively.
[0061] Specifically, as will be described later, the encoding unit 32 generates video packets according to a combination of group IDs t, x, and y. Figure 3 shows packetization when t = 0, x = 0, and y = 0 (see the shaded area in the figure).
[0062] The decoding unit 36 of the receiving-side combining and distributing device 30 selectively receives and decodes video packets using the group ID in accordance with the resolution, image quality, and frame rate of the video to be combined in the video combining unit 33 (scalable decoding and depacketization). This allows the receiving-side combining and distributing device 30 to adaptively receive only the necessary data in accordance with the size of the video to be combined in the video combining unit 33. Note that the decoding unit 36 may be configured to receive and decode video packets selected by the interconnect 50, which serves as a selection mechanism.
[0063] For example, if the receiving combining and distributing device 30 acquires only data (x, y) = (0, 0) without imposing any restrictions on the acquired frame rate, a video of 1 / 16 size will be obtained. Similarly, if the receiving combining and distributing device 30 acquires only data (x, y) = (0, 0), (2, 0), (0, 2), (2, 2), a video of 1 / 4 size will be obtained. Furthermore, if only data (t) = (0) is acquired without imposing any restrictions on (x, y) in each frame, a video of 1 / 4 frame rate will be obtained. Similarly, if only data (t) = (0), (2) is acquired, a video of 1 / 2 frame rate will be obtained.
[0064] If a video packet of the desired scalable code cannot be obtained, the decoding unit 36 can complement it by a method such as enlarging the signal of the group ID of the reference layer (for example, (x, y, z) = (0, 0, 0)). Also, if a video packet of the reference layer cannot be obtained, the decoding unit 36 can complement it with a video packet of another group ID. Also, if a frame packet at a given time cannot be obtained, it can be complemented by substituting data from a frame packet at the previous time.
[0065] Here, the video data is uncompressed or lightly compressed without inter-frame compression, which reduces the delay caused by compression (codec).
[0066] The group ID and its contents t, x, and y can be used as part of an identifier such as the MAC address or IP address of a video packet. This enables selective forwarding using a general L2 (Layer 2) or L3 (Layer 3) switch as the interconnect 50. It is not necessary to use all of t, x, and y. The identifier can also be simplified by setting x = y.
[0067] Furthermore, a priority can be set for video packets for each group ID (or for each combination of group IDs). The priority can be set using CoS (Class of Service), DSCP (Differentiated Services Code Point), IP Precedence values, etc. By setting a high priority for video packets corresponding to a reference layer (e.g., (x, y, z)), reliable video viewing can be achieved.
[0068] Furthermore, when packetizing for each group ID as described above and dividing into multiple packets, a sequence number can be stored in the video packet. This makes it possible to check for missing packets, etc. Furthermore, each video packet can store a frame number to identify the frame.
[0069] According to the above, the video packets that have been packetized in advance for each group ID by the sending-side combining and distribution device 30 can be selectively received and combined by the receiving-side combining and distribution device 30. For example, there is no need to send or receive a basic layer corresponding to low-quality video, and there is no need to compress the data, so delays are reduced.
[0070] [Identifier Coding] Next, coding by the encoding unit 32 will be described with reference to Figures 4 to 8. Figure 4 is a diagram for explaining identification of a group ID using identifier coding.
[0071] 4 shows an example of group IDs and corresponding codes when coding with 4 bits when T, X, or Y is 12, that is, when t, x, or y cycles at 12. Specifically, in this embodiment, group IDs of 0 to 12 are assigned to t, x, or y.
[0072] Each group ID is associated with a 4-bit code. In this embodiment, the rules for incrementing the upper two digits and the lower two digits of the code are different. Specifically, each time the number assigned to the group ID increments by 1, the lower two bits of the code advance in the order 00, 01, 10, and 11, and then cycle through three iterations (a total of three iterations). The receiving combining and distribution device 30 can selectively receive half-size video by identifying only one of the lower two bits of the code (e.g., 0). In other words, in this embodiment, half-size video can be selectively received based on a single matching pattern (e.g., one of the lower two bits = 0). Note that the figure shows an example in which the lowest bit is used for identification. Furthermore, the receiving combining and distribution device 30 can selectively receive quarter-size video by identifying only the lower two bits of the code (e.g., 00). In other words, in this embodiment, quarter-size video can be selectively received based on a single matching pattern (e.g., the lower two bits = 0). This process of associating each group with a code is called identifier coding, and the code used in identifier coding is called a scale code.
[0073] On the other hand, the upper two bits of the code advance in the order 00, 01, 10, and so on, each time the number assigned as the group ID increases by 1, and then cycles around (four times in total). By identifying only the upper two bits of the code (for example, 00), it is possible to selectively receive 1 / 3 size video.
[0074] Furthermore, the group IDs t, x, and y can be interleaved with each other. FIG. 5 is a diagram illustrating bit extension and scale interleaving. FIG. 5 shows an example in which bit extension and interleaving are performed simultaneously (called scale interleaving). As shown in the diagram, this is a method in which the bits of X and Y are packed together into one byte unit for each upper bit and lower bit, and any missing bits are extended by padding with 0s or the like. A code created by this method is called a scale interleaving code. In the example shown in the diagram, 4 bits are padded before the code 1001 corresponding to group ID = 5 for x and the code 0110 corresponding to group ID = a for y, and scale interleaving is performed so that the upper 2 bits and the lower 2 bits are combined. As a result, when the scaling sizes of X and Y are the same and when the scaling sizes are independent scaling bits only, the number of matching tables can be reduced, and packet selection becomes possible using general L2 and L3 switches.
[0075] Here, referring to Figure 6, we will compare the case where the X and Y GIDs are coded as they are with the case where they are coded using a scale code. Unlike the above, if the group IDs are coded using a general binary system, four matching patterns are required, making the selection process more complicated. In other words, if the numbers 0 to 12 are coded in the order 0000, 0001, 0010, 0011, 0100, 0101, 0110, 0111, 1000, 1001, 1010, and 1011, it is not possible to extract four group IDs using a small number of bits, so four matching patterns, for example, 0000, 0011, 0110, and 1001, are required.
[0076] In contrast, in this embodiment, the upper two bits of the code are rotated before reaching 11, so that one matching pattern appears four times corresponding to group IDs from 0 to 12. Therefore, in this embodiment, it is possible to selectively receive 1 / 3 size video based on one matching pattern (for example, upper two bits = 00).
[0077] Specifically, for devices that perform byte (8-bit) matching, the encoding unit 32 can also divide the group ID into bytes. For example, in the case of a code 1001 corresponding to group ID=5 in the figure, the encoding unit 32 can extend the code by padding 6 bits before the most significant 2 bits and 6 bits before the least significant 2 bits, thereby dividing the code into 00000010, 00000001, etc. This makes it possible to reduce the matching table in byte units.
[0078] Next, the advantageous effects of coding using scale codes will be described with reference to Fig. 7. Fig. 7 is a diagram for explaining a comparison between the filter rules when coding the X and Y GIDs as they are and the filter rules when coding using scale codes.
[0079] As shown in FIG. 7A, when X and Y are each reduced to one-third their size, if the GIDs of X and Y are coded as they are, 16 filter rules are required to select 16 group IDs: ((0,0), (3,0), (6,0), (9,0), (0,3), (3,3), (6,3), (9,3), (0,6), (3,6), (6,6), (9,6), (0,9), (3,9), (6,9), (9,9)).
[0080] In contrast, as shown in FIG. 7B, when X and Y are each reduced to 1 / 3 size and scale interleaved using a scale code, at first glance, it may appear that 16 filter rules are required to select 16 group IDs ((0,0), (3,0), (2,0), (1,0), (0,3), (3,3), (2,3), (1,3), (0,2), (3,2), (2,2), (1,2), (0,1), (3,1), (2,1), (1,1)). However, on a byte basis, this can be realized with a filter that matches only the upper bits 00, for example, as shown in FIG. 8. In this case, the lower bits are set as don't care. This allows for a significant reduction in the number of filter rules.
[0081] It should be noted that the scope of the present disclosure is not limited to the above-described identifier coding. Each code may be expressed with a larger or smaller number of bits. In general, by configuring a portion of the code to cycle every time an increase of M occurs (every M-1), it is possible to selectively receive 1 / M size video.
[0082] Next, the specific structure of a video packet will be described with reference to FIG. 9. FIG. 9 is a diagram illustrating the structure of a video packet when a GID is stored in an RTP-extended frame. As described above, the video compositing and distribution system 300 cyclically assigns group IDs of up to T, X, and Y to the video time axis t and screen coordinates x and y (scalable coding). Then, the encoding unit 32 assigns an identifier (code) shown in FIG. 4 for each group ID for each of t, x, and y (identifier coding), and generates a video packet for each group ID (or for each combination of group IDs (x, y, z)) (packetization). In particular, the structure of a video packet when a video packet is generated for each combination of group IDs will be described below.
[0083] As shown in the figure, a code corresponding to each group ID is divided into bytes and stored in an RTP-extended frame 41 in a video packet 40. Specifically, a code corresponding to a time-related group ID is stored as part of a CSRC (Contributing Source) identifier 42a, and a code corresponding to a coordinate-related group ID is stored as part of a CSRC identifier 42b.
[0084] The code corresponding to the time-related group ID is split into two byte-by-byte parts and stored as part of the CSRC identifier 42a (see TimeCode1 and TimeCode2 in the figure). Specifically, each scale code shown in FIG. 4 is expanded by adding 6-bit padding before the most significant two bits and 6-bit padding before the least significant two bits, respectively, and then split. This allows for byte-by-byte reduction of the matching table. For example, when storing the scale code 0011 corresponding to the time-related group ID=3 shown in FIG. 4 in the frame 41 as part of the CSRC identifier 42a, the encoding unit 32 expands the code by adding 6-bit padding before the most significant two bits, 00, and by adding 6-bit padding before the least significant two bits, 11, to split the code into 00000000 (TimeCode1) and 00000011 (TimeCode2).
[0085] Similarly, the code corresponding to the group ID for coordinates x and y is split into two byte units for each of x and y and stored as part of the CSRC identifier 42b (see XCode1, XCode2, YCode1, and YCode2 in the figure). For example, when the code 1001 corresponding to the group ID=5 for x shown in Figure 4 is stored in the frame 42 as part of the CSRC identifier 42b, the encoding unit 32 extends the code by padding 6 bits before the most significant two bits of the code, 10, and by padding 6 bits before the least significant two bits of the code, 01, and splits it into 00000010 (XCode1) and 00000001 (XCode2).
[0086] Also, for example, when the scale code 0110 corresponding to the group ID=a for y shown in FIG. 4 is stored in the frame 42 as part of the CSRC identifier 42b, the encoding unit 32 extends the code by padding 6 bits before the most significant two bits of the code, 01, and by padding 6 bits before the least significant two bits of the code, 10, and divides the code into 00000001 (YCode1) and 00000010 (YCode2).
[0087] The decoding unit 36 in the receiving combining and distributing device 30 identifies only one of the two least significant bits (e.g., 0) for each of TimeCode2, XCode2, and YCode2 and receives the corresponding video packet, allowing the receiving combining and distributing device 30 to selectively receive half-size video. The decoding unit 36 also identifies only the two least significant bits (e.g., 00) for TimeCode2, XCode2, and YCode2 and receives the corresponding video packet, allowing the receiving combining and distributing device 30 to selectively receive quarter-size video.
[0088] Furthermore, the decoding unit 36 identifies only the lowest two bits of TimeCode1, XCode1, and YCode1 and receives the corresponding video packets, allowing the receiving combining and distribution device to selectively receive 1 / 3 size video.
[0089] Furthermore, as described above, 8-bit data mutually interleaved between the t, x, and y codes (see FIG. 5) can be stored as part of the CSRC identifier 42a or CSRC identifier 42b in the frame 42. This allows a reduction in the number of matching tables.
[0090] In this embodiment, 8-bit data is generated by padding 6 bits to the most significant 2 bits and the least significant 2 bits of each code, but the scope of the present disclosure is not limited to this. For example, 4-bit data may be generated by padding 2 bits to the most significant 2 bits and the least significant 2 bits of each code.
[0091] Any of the identifiers constituting the frame 42 may include a sequence number, which makes it possible to check for packet loss, etc. Also, each video packet may have a frame number for identifying the frame.
[0092] The configuration of the video packet is not limited to the above. The group ID can be part of the identifier of the destination MAC address, the source MAC address, or the VLAN ID. Figure 10 shows the configuration of a video packet 50 that stores a GID in the destination MAC address. This makes it possible to perform filtering using an L2 switch.
[0093] When embedding a group ID in a destination MAC address, since the destination is multiple hosts, the I / G (Individual / Group) bit, which is the least significant bit in the first octet, must be set to "1." Therefore, the group ID is embedded in the remaining bits excluding the I / G bit. The figure shows an example in which the group ID is assigned to the fourth through sixth octets. However, the scope of the present disclosure is not limited to this. The group ID can be assigned to other octets, or T, X, and Y can be shortened to fewer than eight bits and combined into fewer octets. In this case, the source MAC address can indicate the identifier of the source composite distribution device, the user data identification child, or both.
[0094] 11 is a diagram illustrating the configuration of a video packet 60 when a GID is stored in a VLAN. When embedding a group ID in a VLAN ID, 12 bits of the VLAN ID can be used. Note that it is possible to use only some of the bits for embedding the group ID, rather than using all 12 bits for embedding the group ID.
[0095] It is also possible to implement the group IDs T, X, and Y separately for the destination MAC address and VLAN ID described above. While implementing a MAC address filtering table is different from implementing a VLAN filtering table, this division makes it less likely that the upper limit on the number of filtering tables in an L2 switch will be exceeded. It is also possible to embed other information, such as a group ID identifier, a source combining / distributing device identifier, or a user data identifier, in any L2 identifier.
[0096] The structure of the video packet is not limited to the above. The group ID can also be implemented using an IP header / UDP header, etc. Fig. 12 shows the structure of a video packet 70 when a GID is stored in the destination IP address. This enables filtering using an L3 switch. For example, the group ID can be part of the identifier of the destination IP address, source IP address, or port number.
[0097] When embedding a group ID in a destination IP address, a multicast address can be used because the destination is multiple hosts. In the case of IPv4, the upper four bits of the first octet are set to 1110, and the group ID is embedded in the remaining bits. The figure shows an example in which the group ID is assigned to the second to fourth octets. However, the scope of the present disclosure is not limited to this, and the group ID can be assigned to other octets, or T, X, and Y can be shortened to fewer than eight bits and combined into fewer octets. In this case, the source IP address can indicate the source synthesis distribution device identifier, the user data identifier, or both.
[0098] In addition, other information such as the group ID identifiers of T, X, and Y, the source combining / distributing device identifier, or the user data identifier can be divided and embedded in any L3 identifier. The group ID and other identifiers set for the above L2 and L3 can be combined or overlapped. By overlapping or dividing the settings, filtering with any parameters can be performed according to the filtering capabilities of L2 and L3 switches with different capabilities, such as filtering table sizes.
[0099] [Effect] When data is compressed and encoded using technologies such as MPEG (Moving Picture Experts Group), delays of several hundred milliseconds or more occur due to the codec. Considering the possibility of remote ensembles being held over video conferences that use this type of screen composition, the delays associated with this distribution significantly impair the feasibility of such a system. For example, in a song with 120 beats per second (120 BPM (beats per minute)), the duration of one beat is 60 / 120 seconds = 500 milliseconds. If this accuracy needs to be matched to 5%, the delay between capturing a video with a camera and displaying it must be reduced to 500 x 0.05 = 25 milliseconds or less.
[0100] In reality, the process from capturing an image with a camera to displaying it requires not only the time required for the compositing process (codec), but also delays due to the image processing time on the camera, the display time on the monitor, and the time required for transmission. However, because devices such as cameras are connected from outside the system, it is difficult to control the processing time required by the device within the system. For this reason, it is possible to reduce delays by shortening the time required for video compositing.
[0101] In this embodiment, the decoding unit 36 of the receiving combining / distributing device 30 can selectively receive video that has been scalably encoded and packetized by the encoding unit 32 of the sending combining / distributing device 30 using uncompressed or light compression. This allows the video combining / distributing system 300, which combines multiple video signals from multiple locations into multiple different video signals for collaborative work with strict low-latency requirements, to reduce the delay, particularly the time from video signal input to video output. Furthermore, by appropriately setting the GID, reduced video can be obtained with low latency and a high degree of flexibility. By dividing X and Y into 12 groups, and using the same selection method for X and Y, it is possible to adjust the size in 12 steps from 1 / 144 size. For sizes in between, video of a similar size is received and spatial interpolation is performed to obtain the desired size with low latency and minimal computation.
[0102] In particular, in this embodiment, the communication band can be reduced to the minimum necessary, which makes it possible to effectively utilize network resources and simplify the network configuration in each device.
[0103] Furthermore, because the code corresponding to the group ID can be used as part of the identifier such as the MAC address or IP address of the video packet, network systems such as general L2 and L3 switches can be used as the interconnect 50 as a data selection mechanism for scalable coding. This eliminates the need to install new servers, etc., and allows the system to be built at low cost.
[0104] In this way, the present disclosure can provide a video compositing and distribution system 300 that can reduce communication costs for uncompressed video or lightly compressed video.
[0105] The device of the present invention can also be realized by a computer and a program, and the program can be recorded on a recording medium or provided via a network. The program of the present disclosure is a program for causing a computer to realize each function of the device according to the present disclosure, and a program for causing a computer to execute each procedure of the method executed by the device according to the present disclosure.
[0106] 10: Control device 30, 30A, 30B, 30C, 30D: Composing and distribution device 31: Video input IF 32: Encoding unit 33: Video composition unit 34: Video output IF 35: Network IF 36: Decoding unit 37: Video request unit 38: Network control unit 39: Control reception unit 40: Video packet 41: Frame 42a, 42b: CSRC identifier 50: Interconnect 300: Video composition and distribution system
Claims
1. A video compositing and distribution system comprising: a sending-side compositing and distribution device that selects multiple partial images with different coordinates from a single image, combines the multiple partial images, and transmits them; and a receiving-side compositing and distribution device that decodes the image using the data of the multiple combined partial images transmitted from the sending-side compositing and distribution device.
2. The video compositing and distribution system according to claim 1, wherein the sending-side compositing and distribution device assigns a common identifier to the plurality of partial images to be selected, and selects a plurality of different partial images from the single image based on the identifier, and the receiving-side compositing and distribution device receives the data of the plurality of partial images based on the identifier and decodes the images.
3. The video composition distribution system according to claim 2, wherein the identifier rotates on the single image at a predetermined cycle.
4. The video compositing and distribution system of claim 3, wherein the images are frame images included in the video, the sending-side compositing and distribution device assigns a common identifier to frame images that rotate at a predetermined time period among multiple frame images included in the video at different times, and the receiving-side compositing and distribution device receives the frame images based on the identifier.
5. The video compositing and distribution system according to claim 2, wherein the identifier has a bit length of two or more, the sending-side combining and distribution device divides the identifier into two or more parts and stores each divided identifier in a predetermined packet area, and the receiving-side combining and distribution device obtains the identifier by reading the predetermined packet area.
6. A video compositing and distribution system according to any one of claims 2 to 5, wherein the sending-side combining and distribution device generates a packet for each of the same groups, encodes the identifier, stores it in a predetermined area of the packet, and selectively transmits the packet in response to a transmission request based on the identifier from the receiving-side combining and distribution device; and the receiving-side combining and distribution device selectively receives the packet by identifying a part of the encoded identifier stored in the predetermined area, and decodes the image using one or more of the packets.
7. The video compositing and distribution system according to claim 6, further comprising an interconnect interposed between the sending-side compositing and distribution device and the receiving-side compositing and distribution device, which receives the packet having the encoded identifier stored in the header portion from the receiving-side compositing and distribution device in response to a transmission request based on the identifier of the receiving-side compositing and distribution device, and selectively transfers the packet to the sending-side compositing and distribution device.
8. A video synthesis distribution method comprising the steps of: selecting multiple partial images with different coordinates from a single image, synthesizing and transmitting the multiple partial images; and decoding the image using the data of the multiple synthesized partial images.
Citation Information
Patent Citations
Transmission apparatus, transmission method, reception apparatus, reception method and signal transmission system
JP2011172164A
Signal processor, signal processing method, program, and signal transmission system
JP2015076704A
Image output device, image processing device, and display device
JP2018196039A
Processing device and control method thereof
JP2020120212A
Method for detecting 2si mapping errors in 4k / 8k video and correction processing method and device therefor
JP2021089335A