Synthetic image generation apparatus and synthetic image generation method
The data generation device improves the combination and synchronization of images from multiple viewpoints by using positional and directional information to generate wide-area top-view images, addressing the limitations of existing systems and enhancing convoy driving assistance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-10
AI Technical Summary
Existing data generation devices for generating top-view images from multiple captured images lack the capability to accurately combine and synchronize the images from different viewpoints, limiting their effectiveness in providing comprehensive environmental awareness, especially in convoy driving scenarios.
A data generation device that acquires images from multiple cameras on moving objects, determines the composite position and orientation of these images using positional and directional information, and generates a wide-area top-view image by mapping these images onto a virtual space, synchronizing them based on time and positional relationships.
Enables accurate and synchronized generation of wide-area top-view images that provide a comprehensive understanding of the environment, assisting drivers in convoy driving by overcoming obstructions and enhancing situational awareness.
Smart Images

Figure 2026062899000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a data generation device and a data generation method.
Background Art
[0002] Conventionally, a data generation device (i.e., an image generation device) that generates a top-view image from a plurality of captured images has been proposed (see, for example, Patent Document 1). Such a top-view image is an image of the surroundings including a vehicle as seen from above the vehicle, and is used, for example, for assistance during parking.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In such a data generation device, further improvement is required.
[0005] Therefore, the present disclosure provides a data generation device capable of achieving further improvement.
Means for Solving the Problems
[0006] A composite image generation apparatus according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, acquires a first image generated using a camera provided on a first mobile body, acquires a second image generated using a camera provided on a second mobile body, generates a composite image from the first image and the second image, determines the composite position of the first image and the second image based on the positional information of the first and second mobile bodies and the matching result of feature points extracted from the first image and feature points extracted from the second image, wherein the first image and the second image are, respectively, a part of the top view image of the first mobile body and a part of the top view image of the second mobile body, the top view image is an image of the surroundings including the mobile body corresponding to the top view image viewed from above the mobile body, and the part of the top view image is a portion specified according to the relative positional relationship between the first mobile body and the second mobile body.
[0007] A composite image generation apparatus according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, acquires a first image generated using a camera provided on a first mobile body, acquires a second image generated using a camera provided on a second mobile body, generates a composite image from the first image and the second image, and in generating the composite image, determines the composite position of the first image and the second image based on the positional information of the first and second mobile bodies and the matching result of feature points extracted from the first image and feature points extracted from the second image.
[0008] These general or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium. [Effects of the Invention]
[0009] This disclosure provides a data generation device that can achieve further improvements. [Brief explanation of the drawing]
[0010] [Figure 1] Figure 1 is a block diagram showing the functional configuration of the data generation device in the embodiment. [Figure 2] Figure 2 is a block diagram showing the detailed functional configuration of the data generation device in the embodiment. [Figure 3] Figure 3 shows an example of a convoy of three vehicles in the embodiment. [Figure 4] Figure 4 shows an example of how a wide-area top-view image is generated from top-view images of each vehicle in the embodiment. [Figure 5A] Figure 5A is a flowchart showing the overall processing operation of the data generation device in the embodiment. [Figure 5B] Figure 5B is a flowchart showing the synthesis position determination process by the wide-area synthesis unit in the embodiment. [Figure 6A] Figure 6A is a block diagram showing an example of the implementation of the data generation device in the embodiment. [Figure 6B] Figure 6B is a flowchart showing the processing operation of a data generation device equipped with a circuit and memory in an embodiment. [Figure 7] Figure 7 is an overall diagram of the content supply system that realizes the content distribution service. [Figure 8] Figure 8 shows an example of an encoding structure during scalable encoding. [Figure 9] Figure 9 shows an example of an encoding structure during scalable encoding. [Figure 10] Figure 10 shows an example of how a web page is displayed. [Figure 11] Figure 11 shows an example of how a web page is displayed. [Figure 12] Figure 12 shows an example of a smartphone. [Figure 13] Figure 13 is a block diagram showing an example of a smartphone configuration.
Mode for Carrying Out the Invention
[0011] A data generation device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit. In operation, the circuit acquires sensing data configured based on the sensing results of a plurality of sensors provided in each of a plurality of moving objects from each of the plurality of moving objects, generates synthetic data by mapping the sensing data of each of the plurality of moving objects onto a virtual space, and in the generation of the synthetic data, determines the position of the sensing data mapped onto the virtual space according to at least the position in the real space of the moving object corresponding to the sensing data. For example, the moving object is a vehicle, the sensing data is a top view image, and the synthetic data is a wide-area top view image.
[0012] Thereby, not only the environment around one moving object but also the environments including the surroundings of each of the plurality of moving objects are sensed, and synthetic data showing the sensed environments on the virtual space is generated, so that a wider range of environments can be appropriately grasped. Further, when each of the plurality of moving objects moves, the position of the sensing data of the moving object on the virtual space can be changed according to the position of the moving object in the real space after the movement. Therefore, even when the plurality of moving objects move, the synthetic data can be made to follow their movements.
[0013] Each of the plurality of sensors may be a camera, and the circuit may acquire an image as the sensing data.
[0014] Thereby, synthetic data showing the environments including the surroundings of each of the plurality of moving objects as images is generated, so that by looking at the images, a wide range of environments can be easily grasped visually.
[0015] Further, in determining the position of the sensing data, the circuit may extract feature points from an image that is the sensing data, and determine the position of the sensing data according to the extracted feature points and the position of the moving body in the real space corresponding to the sensing data.
[0016] Thereby, not only the positions of the respective moving bodies but also the positions of the feature points of the image are used to determine the position of the sensing data, so that the sensing data can be mapped more accurately.
[0017] Further, the circuit further obtains position information indicating the position of each of the plurality of moving bodies in the real space at the time when the sensing data of the moving body is generated from each of the plurality of moving bodies, and in generating the composite data, determines the position of the sensing data obtained from the moving body in the virtual space based on the positions indicated by the respective position information of the plurality of moving bodies.
[0018] Thereby, since position information is obtained from each of the plurality of moving bodies, the positions of those moving bodies in the real space can be easily specified, and the processing burden of determining the position of the sensing data can be reduced.
[0019] Further, the circuit further obtains direction information indicating the traveling direction of each of the plurality of moving bodies at the time when the sensing data of the moving body is generated from each of the plurality of moving bodies, and in generating the composite data, determines the orientation of the sensing data obtained from the moving body in the virtual space based on the traveling directions indicated by the respective direction information of the plurality of moving bodies.
[0020] Thereby, based on the direction information of the moving body, the orientation of the sensing data of the moving body in the virtual space is determined, so that the sensing data can be mapped in an appropriate orientation. As a result, the sensing data can be mapped more accurately.
[0021] Furthermore, the circuit may, in acquiring the sensing data, periodically acquire the sensing data and time information indicating the time when the sensing data was generated from each of the plurality of moving bodies, and in generating the composite data, for each of the plurality of moving bodies, select a specific sensing data from the plurality of sensing data periodically acquired from that moving body, the time indicated by the time information corresponding to the sensing data being within a predetermined period, and map the selected specific sensing data onto the virtual space.
[0022] As a result, the multiple sensing data mapped onto the virtual space are specific sensing data generated within a predetermined period. Therefore, it is possible to properly synchronize the multiple sensing data mapped onto the virtual space.
[0023] Furthermore, the circuit may acquire the sensing data from each of the multiple moving objects that are in a predetermined positional relationship.
[0024] For example, in a predetermined positional relationship, multiple moving objects travel in a convoy lined up in a row. In such cases, composite data is generated, allowing one of the multiple moving objects to easily understand the environment surrounding other moving objects in front of or behind it.
[0025] Furthermore, each of the multiple moving bodies is a vehicle, and in the predetermined positional relationship, the multiple moving bodies move in a line, and in the generation of the composite data, the circuit generates a wide-area top-view image, which is the composite data, by mapping the top-view images, which are the sensing data acquired from each of the multiple moving bodies lined up in a line, onto the two-dimensional space, which is the virtual space, and the top-view image may be an image of the surroundings including the moving body corresponding to the top-view image, viewed from above the moving body.
[0026] This generates a wide-area top-view image, which is an image of multiple vehicles traveling in a convoy viewed from above. Therefore, even if the view of one of the vehicles in the convoy is obstructed by other vehicles in front of or behind it, the wide-area top-view image allows the driver to easily perceive the environment, including the surroundings of those vehicles. As a result, appropriate assistance can be provided for driving in a convoy.
[0027] Furthermore, the circuit may display the image represented by the synthesized data on a display.
[0028] This allows, for example, the driver of a vehicle to be shown synthesized data, enabling appropriate assistance in driving that vehicle.
[0029] The embodiments will be described in detail below with reference to the drawings.
[0030] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit the scope of the claims. Furthermore, among the components in the following embodiments, those not described in the independent claim representing the highest-level concept will be described as optional components.
[0031] (Embodiment 1) Figure 1 is a block diagram showing the functional configuration of the data generation device in this embodiment.
[0032] The vehicle 100 includes a first sensor 101, a second sensor 102, a third sensor 103, and a fourth sensor 104, a position and direction detection unit 110, and a data generation device 200.
[0033] Each of the first sensor 101, second sensor 102, third sensor 103, and fourth sensor 104 is configured as, for example, a camera. Specifically, the first sensor 101 is a camera that photographs the front of the vehicle 100. The second sensor 102 is a camera that photographs the left side of the vehicle 100. The third sensor 103 is a camera that photographs the right side of the vehicle 100. The fourth sensor 104 is a camera that photographs the rear of the vehicle 100.
[0034] The position and direction detection unit 110 detects the current position and direction of travel of the vehicle 100 and outputs position information indicating the detected position and direction information indicating the detected direction of travel to the data generation device 200. For example, the position and direction detection unit 110 periodically performs detection and outputs the position information and direction information, which are the detection results, to the data generation device 200. Specifically, the position and direction detection unit 110 detects the position and direction of travel using GNSS (Global Navigation Satellite System). In other words, the position and direction detection unit 110 detects the position and direction of travel by receiving signals transmitted from satellites.
[0035] The data generation device 200 generates a wide-area top-view image as composite data by combining a top-view image of the vehicle 100 with a top-view image of at least one surrounding vehicle. Such a data generation device 200 includes an acquisition unit 210, a synthesis unit 220, a display unit 230, and an output unit 240.
[0036] The acquisition unit 210 acquires a top-view image from each of at least one surrounding vehicle. The top-view image acquired from this surrounding vehicle is an image of the surrounding area, including the surrounding vehicle, as seen from above the surrounding vehicle.
[0037] The image synthesis unit 220 acquires captured images from each of the first sensor 101, second sensor 102, third sensor 103, and fourth sensor 104. The image synthesis unit 220 then synthesizes the captured images acquired from these sensors to generate a top-view image of the vehicle 100. This top-view image is an image of the area including the vehicle 100 as seen from above the vehicle 100.
[0038] Furthermore, the synthesis unit 220 generates a wide-area top-view image by combining the top-view image of the vehicle 100 with at least one top-view image acquired by the acquisition unit 210. When generating the wide-area top-view image, the synthesis unit 220 uses the position information and direction information output from the position and direction detection unit 110.
[0039] The display unit 230 consists of a liquid crystal display, a plasma display, or an organic EL (Electro-Luminescence) display, and displays a wide-area top-view image generated by the synthesis unit 220.
[0040] The output unit 240 transmits the top-view image of the vehicle 100 generated by the synthesis unit 220 to at least one surrounding vehicle.
[0041] Here, the at least one surrounding vehicle that transmits a top-view image to the data generation device 200 of vehicle 100 may be the same as or different from the at least one surrounding vehicle that receives the top-view image of vehicle 100. Furthermore, each of the at least one surrounding vehicle that transmits a top-view image to the data generation device 200 of vehicle 100 may be equipped with multiple sensors, similar to vehicle 100, and may generate a top-view image of that surrounding vehicle based on the sensing results of those sensors.
[0042] Therefore, the data generation device 200 in this embodiment acquires sensing data from each of the multiple moving objects, based on the sensing results of each of the multiple sensors provided on that moving object. The multiple moving objects consist of a vehicle 100 and at least one surrounding vehicle. The sensing data is a top-view image. The data generation device 200 then generates composite data, which is a wide-area top-view image, by mapping the sensing data of each of the multiple moving objects onto a virtual space. In this embodiment, when generating the composite data, the data generation device 200 determines the position of the sensing data mapped onto the virtual space according to at least the real-space position of the moving object corresponding to that sensing data. The real-space position of the moving object is, for example, the position of the vehicle 100 or surrounding vehicle detected by GNSS or the like.
[0043] This allows for sensing not only the environment surrounding vehicle 100, but also the environment surrounding each of the multiple vehicles, and generates composite data that represents the sensed environment in virtual space, thus enabling a more accurate understanding of a wider range of environments. Furthermore, when each of the multiple vehicles moves, the position of the vehicle's sensing data in virtual space can be changed according to the vehicle's position in real space after its movement. Therefore, even when multiple vehicles move, the composite data can be made to follow their movements.
[0044] Furthermore, in this embodiment, each of the multiple sensors provided in each of the multiple vehicles, such as the first sensor 101, the second sensor 102, the third sensor 103, and the fourth sensor 104, is a camera. The data generation device 200 in this embodiment acquires images as sensing data. As a result, composite data is generated that shows the environment including the surroundings of each of the multiple vehicles as an image, so that by looking at that image (i.e., a wide-area top-view image), a wide area of the environment can be easily grasped visually.
[0045] Furthermore, in this embodiment, the image represented by the composite data is displayed on the display unit 230. This allows, for example, the driver of the vehicle 100 to see the composite data, thereby providing appropriate support for driving the vehicle 100.
[0046] Figure 2 is a block diagram showing the detailed functional configuration of the data generation device 200.
[0047] The acquisition unit 210 of the data generation device 200 includes a first communication unit 211, a reception control unit 212, a data reception unit 213, and a first format conversion unit 214.
[0048] The reception control unit 212 controls the first communication unit 211, the data reception unit 213, and the first format conversion unit 214.
[0049] The first communication unit 211 establishes a communication path with a surrounding vehicle based on control by the receiving control unit 212 and requests data transmission from that surrounding vehicle. The surrounding vehicle from which data transmission is requested will be referred to below as the requesting surrounding vehicle. At this time, the first communication unit 211 exchanges information with the requesting surrounding vehicle regarding the data formats that each vehicle can handle, based on control by the receiving control unit 212.
[0050] The data receiving unit 213 receives a top-view image of a nearby vehicle from the requesting vehicle using the communication channel established by the first communication unit 211. This top-view image is accompanied by control information including time information, position information, and direction information. The time information indicates the time when the top-view image of the nearby vehicle was generated. The position information indicates the position of the nearby vehicle at the time the top-view image was generated, and the direction information indicates the direction of travel of the nearby vehicle at the time the top-view image was generated.
[0051] The first format conversion unit 214 converts the data format of the top-view image received by the data reception unit 213 into a predetermined data format and outputs it to the synthesis unit 220.
[0052] Here, when a nearby vehicle receives a data transmission request from the acquisition unit 210, it periodically transmits the top-view images described above. In other words, the acquisition unit 210 acquires an image signal containing multiple top-view images generated by the nearby vehicle and outputs the format-converted image signal to the synthesis unit 220. The control information described above is added to the multiple top-view images acquired in this way.
[0053] Furthermore, if the image signal is encoded, the first format conversion unit 214 may decode the encoded image signal. In other words, the first format conversion unit 214 may be configured as a decoding device, an image decoding device, or a video decoding device. For example, the first format conversion unit 214 decodes an image signal encoded based on a video compression standard such as H.264 or HEVC (High Efficiency Video Coding) according to a decoding method, image decoding method, or video decoding method corresponding to that standard.
[0054] The data generation device 200's synthesis unit 220 includes an image storage unit 221, a wide-area synthesis unit 222, and a top-view image generation unit 223.
[0055] The top-view image generation unit 223 acquires captured images from each of the first sensor 101, second sensor 102, third sensor 103, and fourth sensor 104. The top-view image generation unit 223 then generates a top-view image of the vehicle 100 by combining the captured images acquired from these sensors. Specifically, the top-view image generation unit 223 periodically generates top-view images. That is, each of the first sensor 101, second sensor 102, third sensor 103, and fourth sensor 104 takes pictures at a predetermined frame rate and repeatedly outputs captured images according to that frame rate. The top-view image generation unit 223 generates a top-view image of the vehicle 100 by combining the captured images taken by these sensors at substantially the same timing. For example, the top-view image generation unit 223 repeatedly generates top-view images according to the above frame rate and outputs them to the wide-area synthesis unit 222.
[0056] Here, the top-view image generation unit 223 adds control information, including time information, position information, and direction information, to the generated top-view image each time a top-view image is generated. The time information indicates the time when the top-view image was generated. The position information and direction information indicate the position and direction detected by the position and direction detection unit 110. Specifically, the position information indicates the position of the vehicle 100 at the time the top-view image was generated, and the direction information indicates the direction of travel of the vehicle 100 at that time.
[0057] The wide-area synthesis unit 222 acquires image signals of surrounding vehicles from the acquisition unit 210 and periodically acquires top-view images of vehicle 100 from the top-view image generation unit 223. Each time the wide-area synthesis unit 222 acquires a top-view image of vehicle 100 from the top-view image generation unit 223, it selects a top-view image of a surrounding vehicle corresponding to that image from the image signals. For example, the wide-area synthesis unit 222 selects a top-view image of a surrounding vehicle using the time information attached to the top-view image of vehicle 100 and the time information attached to each top-view image included in the image signals of the surrounding vehicles. Specifically, the wide-area synthesis unit 222 selects a top-view image of a surrounding vehicle that has time information attached indicating the same time as the time indicated by the time information of the top-view image of vehicle 100, or a time within a predetermined error range.
[0058] The wide-area synthesis unit 222 then generates a wide-area top-view image by combining the top-view image of the vehicle 100 with the top-view images of selected surrounding vehicles, and stores it in the image storage unit 221.
[0059] In this embodiment, the wide-area synthesis unit 222 periodically acquires top-view images, which are sensing data, and time information indicating the time when the top-view image was generated, from each of the multiple vehicles, including the vehicle 100 and surrounding vehicles. The wide-area synthesis unit 222 then selects a specific top-view image from the multiple top-view images periodically acquired from each of the multiple vehicles, the time indicated by the time information corresponding to that top-view image falls within a predetermined period. The wide-area synthesis unit 222 then maps the selected specific top-view image onto a virtual space.
[0060] As a result, the multiple top-down view images mapped onto the virtual space are specific top-down view images generated within a predetermined period. Therefore, it is possible to properly synchronize the multiple top-down view images mapped onto the virtual space.
[0061] The image storage unit 221 is a recording medium for storing the wide-area top-view image generated by the wide-area synthesis unit 222. For example, the image storage unit 221 may be a hard disk, RAM (Read Only Memory), ROM (Random Access Memory), or semiconductor memory. Such an image storage unit 221 may be volatile or non-volatile.
[0062] The output unit 240 of the data generation device 200 includes a second communication unit 241, a transmission control unit 242, a data transmission unit 243, and a second format conversion unit 244.
[0063] The transmission control unit 242 controls the second communication unit 241, the data transmission unit 243, and the second format conversion unit 244.
[0064] When the second communication unit 241 receives a data transmission request from a nearby vehicle, it establishes a communication channel with that nearby vehicle based on control by the transmission control unit 242. The nearby vehicle that made the data transmission request will be referred to below as the requesting nearby vehicle. Furthermore, the second communication unit 241 exchanges information with the requesting nearby vehicle regarding the data formats that each vehicle can handle, based on control by the transmission control unit 242.
[0065] The second format conversion unit 244 acquires the generated top-view image of the vehicle 100 each time the top-view image generation unit 223 generates the top-view image of the vehicle 100. Then, based on the control from the transmission control unit 242, the second format conversion unit 244 converts the data format of the top-view image to the data format corresponding to the requesting surrounding vehicle.
[0066] Furthermore, the second format conversion unit 244 may encode at least one top-view image of the vehicle 100. In other words, the second format conversion unit 244 may be configured as an encoding device, an image encoding device, or a video encoding device. For example, the second format conversion unit 244 encodes these top-view images according to an encoding method, image encoding method, or video encoding method based on a video compression standard such as H.264 or HEVC (High Efficiency Video Coding).
[0067] The data transmission unit 243 acquires the converted top-view image each time the data format of the top-view image of the vehicle 100 is converted by the second format conversion unit 244. Then, the data transmission unit 243 transmits the top-view image of the vehicle 100 to the requesting surrounding vehicle using the communication channel established by the second communication unit 241. In other words, the data transmission unit 243 transmits an image signal containing multiple top-view images of the vehicle 100 to the requesting surrounding vehicle. The control information described above is added to each top-view image included in this image signal.
[0068] Furthermore, if at least one top-view image of the vehicle 100 is encoded by the second format conversion unit 244, the data transmission unit 243 transmits the stream or bitstream generated by that encoding to the requesting nearby vehicle.
[0069] Thus, when encoding based on the above-mentioned video compression standard is performed on the top-view image, the amount of data in the top-view image can be reduced, and processing delays can be suppressed. For example, the display delay of a wide-area top-view image constructed using that top-view image can be suppressed.
[0070] Figure 3 shows an example of a convoy of three vehicles.
[0071] For example, as shown in Figure 3, vehicles C1, C2, and C3 are traveling in a convoy. That is, vehicles C1, C2, and C3 are arranged in a line and traveling on the road in the same direction. For example, each of vehicles C1, C2, and C3 has the same configuration as vehicle 100 described above.
[0072] In this case, vehicle C1 generates a top-view image of vehicle C1, vehicle C2 generates a top-view image of vehicle C2, and vehicle C3 generates a top-view image of vehicle C3. The top-view image of vehicle C1 is an image of the area including vehicle C1 as seen from above vehicle C1, as shown in Figure 3. The top-view images of vehicles C2 and C3 are also images seen from above their respective vehicles, similar to the top-view image of vehicle C1.
[0073] Furthermore, the top-view image of vehicle C1 is accompanied by control information indicating the time, position, and direction of travel at the time the top-view image was generated. Similarly, the top-view image of vehicle C2 is accompanied by control information indicating the time, position, and direction of travel at the time the top-view image was generated, and the top-view image of vehicle C3 is accompanied by control information indicating the time, position, and direction of travel at the time the top-view image was generated.
[0074] For example, let's say vehicle C2 is the local vehicle, and vehicles C1 and C3 are each surrounding vehicles of vehicle C2. In this case, vehicle C2, being the local vehicle, receives a top-view image of vehicle C1 from the surrounding vehicle C1, and receives a top-view image of vehicle C3 from the other surrounding vehicle C3. Furthermore, vehicle C2 generates a top-view image of vehicle C2, and generates a wide-area top-view image by combining its own top-view image with the top-view images of the two surrounding vehicles.
[0075] Furthermore, when vehicle C2 receives a data transmission request from each of the two surrounding vehicles, it transmits a top-view image of vehicle C2 to those surrounding vehicles.
[0076] Figure 4 shows an example of how a wide-area top-view image is generated from top-view images of each vehicle.
[0077] For example, vehicle C2, the vehicle itself, composites a top-view image of vehicle C2 onto a lane image based on map information. The map information is, for example, information used in a car navigation system and is stored on the recording medium of vehicle C2. Alternatively, the map information may be acquired by vehicle C2 via a network such as the internet and stored on the recording medium of vehicle C2.
[0078] Specifically, the wide-area synthesis unit 222, provided in the data generation device 200 for vehicle C2, acquires a top-view image of vehicle C2 from the top-view image generation unit 223. The wide-area synthesis unit 222 then identifies the position indicated by the control information attached to the top-view image of vehicle C2. The wide-area synthesis unit 222 then extracts the lane image associated with that identified position from the map information. For example, the identified position is the center of the lane image.
[0079] Next, the wide-area synthesis unit 222 identifies the direction of travel indicated by the control information of the top-view image of vehicle C2. Based on the identified direction of travel, the wide-area synthesis unit 222 changes the orientation of the top-view image of vehicle C2 and superimposes the top-view image onto the lane image at the previously identified position. For example, the wide-area synthesis unit 222 rotates the top-view image and superimposes it onto the lane image so that the orientation of the top-view image matches the orientation of the lane image. This generates a provisional wide-area top-view image.
[0080] Next, the wide-area synthesis unit 222 generates a final wide-area top view image by combining the top view image of vehicle C1, which is a surrounding vehicle, and the top view image of another surrounding vehicle, vehicle C3, with its provisional wide-area top view image.
[0081] Here, when the wide-area synthesis unit 222 synthesizes the top-view images of surrounding vehicles onto a temporary wide-area top-view image, it performs a synthesis position determination process to determine the synthesis position and orientation of the top-view image. The wide-area synthesis unit 222 changes the orientation of the top-view images of the surrounding vehicles to the orientation determined by the synthesis position determination process. Then, the wide-area synthesis unit 222 superimposes the top-view images of the surrounding vehicles onto the synthesis position in the temporary wide-area top-view image, that is, the synthesis position determined by the synthesis position determination process.
[0082] Figure 5A is a flowchart showing the overall processing operation of the data generation device 200.
[0083] First, the synthesis unit 220 of the data generation device 200 generates a top-view image of the vehicle 100, which is the vehicle itself, and treats this top-view image as a temporary wide-area top-view image (step S10). At this time, the synthesis unit 220 may also generate a temporary top-view image by superimposing the top-view image of the vehicle 100 onto the lane image, as described above.
[0084] Next, the data generation device 200 extracts at least one feature point from the provisional wide-area top-view image (step S20). For example, the feature point is obtained by image processing such as SIFT (Scale-invariant feature transform), SURF (Speed-Upped Robust Feature), ORB (Oriented-BRIEF), or AKAZE (Accelerated KAZE).
[0085] Then, the data generation device 200 performs the processing in steps S30 to S50 for each of at least one surrounding vehicle.
[0086] In step S30, the wide-area synthesis unit 222 acquires top-view images of surrounding vehicles from image signals transmitted from surrounding vehicles, corresponding to the time information of the top-view image of vehicle 100 generated in step S10. Furthermore, the data generation device 200 acquires direction information and position information attached to the top-view images of the surrounding vehicles.
[0087] In step S40, the wide-area synthesis unit 222 uses at least one feature point extracted in step S20 to perform a synthesis position determination process that determines the synthesis position of the top-view images of surrounding vehicles in the provisional wide-area top-view image.
[0088] In step S50, the wide-area synthesis unit 222 synthesizes the top-view images of surrounding vehicles at the synthesis position of the provisional wide-area top-view image determined by the synthesis position determination process.
[0089] The final wide-area top-view image is generated by performing these steps S30 to S50 for each of at least one surrounding vehicle.
[0090] The data generation device 200 then displays the generated wide-area top-view image on the display unit 230, thereby presenting the wide-area top-view image to the driver of the vehicle (step S60).
[0091] Figure 5B is a flowchart showing the synthesis position determination process by the wide-area synthesis unit 222. In other words, Figure 5B is a flowchart that shows in detail the process of step S40 in Figure 5A.
[0092] The wide-area synthesis unit 222 rotates the top-view images of surrounding vehicles according to their relative directions of travel to the vehicle itself (step S41). For example, the wide-area synthesis unit 222 identifies the direction of travel of the vehicle itself, indicated by the direction information attached to the top-view image of the vehicle itself, and the direction of travel of the surrounding vehicles, indicated by the direction information attached to the top-view images of the surrounding vehicles. The wide-area synthesis unit 222 then rotates the top-view images of the surrounding vehicles by the difference in their directions of travel.
[0093] Next, the wide-area synthesis unit 222 determines candidate synthesis positions in the provisional wide-area top-view image based on the relative positions of surrounding vehicles with respect to its own vehicle (step S42). For example, the wide-area synthesis unit 222 identifies the position of its own vehicle, indicated by the position information attached to the top-view image of its own vehicle, and the positions of the surrounding vehicles, indicated by the position information attached to the top-view images of the surrounding vehicles. Then, the wide-area synthesis unit 222 determines candidate synthesis positions in the provisional wide-area top-view image based on the relative relationship of these positions.
[0094] Furthermore, the wide-area synthesis unit 222 extracts at least one feature point from the top-view image of the surrounding vehicles (step S43).
[0095] Then, the wide-area synthesis unit 222 matches the feature points of the provisional wide-area top-view image extracted in step S20 shown in Figure 5A with the feature points of the top-view image of the surrounding vehicles extracted in step S43. As a result, the wide-area synthesis unit 222 refines the candidate synthesis position determined in step S42 (step S44).
[0096] As described above, the wide-area synthesis unit 222 in this embodiment extracts feature points from the sensing data image (i.e., top-view image), and determines the synthesis position of the top-view image according to the extracted feature points and the real-space position of the vehicle corresponding to the top-view image. As a result, the position of the top-view image is determined not only based on the position of each vehicle but also on the feature points of the image, so the top-view image can be mapped more accurately.
[0097] Furthermore, in this embodiment, the synthesis unit 220 acquires position information from each of the multiple vehicles, indicating the real-world position of that vehicle at the time the top-view image of that vehicle is generated. The wide-area synthesis unit 222 then determines the virtual-space position of the top-view image acquired from the vehicle based on the position indicated by the position information of each of the multiple vehicles. As a result, since position information is acquired from each of the multiple vehicles, the real-world positions of those vehicles can be easily identified, and the processing burden of determining the position of the top-view image can be reduced.
[0098] Furthermore, in this embodiment, the synthesis unit 220 acquires direction information from each of the multiple vehicles, indicating the direction of travel of that vehicle at the time the top-view image of that vehicle is generated. The wide-area synthesis unit 222 then determines the orientation of the top-view image acquired from the vehicle in virtual space based on the direction of travel indicated by the direction information of each of the multiple vehicles. As a result, the orientation of the top-view image of a vehicle in virtual space is determined based on the vehicle's direction information, allowing the top-view image to be mapped in the appropriate orientation. Consequently, the top-view image can be mapped more accurately.
[0099] Furthermore, the data generation device 200 in this embodiment acquires top-view images from each of a plurality of vehicles in a predetermined positional relationship. For example, in a predetermined positional relationship, multiple moving objects are traveling in a convoy in a line. In such a case, for example, a wide-area top-view image is generated as composite data, so that one of the multiple vehicles can easily grasp the surrounding environment of the other vehicles in front of or behind it.
[0100] Furthermore, the wide-area synthesis unit 222 in this embodiment generates a wide-area top-view image, which is synthesized data, by mapping top-view images acquired from each of the multiple vehicles lined up in a row onto a two-dimensional virtual space. This top-view image is an image of the surroundings, including the vehicle corresponding to that top-view image, as seen from above the vehicle.
[0101] This generates a wide-area top-view image, which is an image of multiple vehicles traveling in a convoy viewed from above. Therefore, even if the view of one of the vehicles in the convoy is obstructed by other vehicles in front of or behind it, the wide-area top-view image allows the driver to easily perceive the environment, including the surroundings of those vehicles. As a result, appropriate assistance can be provided for driving in a convoy.
[0102] Figure 6A is a block diagram showing an example of the implementation of the data generation device 200 in this embodiment. The data generation device 200 includes a circuit 201 and a memory 202. For example, the multiple components of the data generation device 200 shown in Figures 1 and 2 are implemented by the circuit 201 and memory 202 shown in Figure 6A.
[0103] Circuit 201 is an information processing circuit and is connected to memory 202. For example, circuit 201 is a dedicated or general-purpose electronic circuit that generates data such as wide-area top-view images. Circuit 201 may also be a processor such as a CPU. Alternatively, circuit 201 may be a collection of multiple electronic circuits. Furthermore, for example, circuit 201 may play the role of multiple components of the data generation device 200 shown in Figures 1 and 2, excluding the component for storing information.
[0104] Memory 202 is a general-purpose or dedicated memory that stores information for circuit 201 to generate data such as wide-area top-view images. Memory 202 may be an electronic circuit. Memory 202 may also be included in circuit 201. Memory 202 may also be a collection of multiple electronic circuits. Memory 202 may also be a magnetic disk or an optical disk, or it may be described as storage or a recording medium. Memory 202 may also be a non-volatile memory or a volatile memory.
[0105] For example, memory 202 may store images for generating a wide-area top-view image, or it may store a program for circuit 201 to generate a wide-area top-view image.
[0106] Furthermore, for example, the memory 202 may play the role of an information storage component among the multiple components of the data generation device 200 shown in Figures 1 and 2, respectively. Specifically, the memory 202 may play the role of the image storage unit 221 shown in Figure 2.
[0107] Furthermore, it is not necessary for the data generation device 200 to implement all of the components shown in Figures 1 and 2, nor is it necessary for all of the processes described above to be performed. Some of the components shown in Figures 1 and 2 may be included in other devices, and some of the processes described above may be performed by other devices.
[0108] Figure 6B is a flowchart showing the processing operation of the data generation device 200, which includes circuit 201 and memory 202.
[0109] In operation, the circuit 201 connected to the memory 202 acquires sensing data from each of the multiple moving objects, based on the sensing results of each of the multiple sensors provided on that moving object (step S1). Next, the circuit 201 generates composite data by mapping the sensing data of each of the multiple moving objects onto a virtual space (step S2). Here, in generating the composite data, the circuit 201 determines the position of the sensing data mapped onto the virtual space according to the real-space position of the moving object corresponding to that sensing data. The sensing data acquired by the circuit 201 from each of the multiple moving objects may be the top-view image described above, or it may be the sensing results (i.e., captured images) of each of the multiple sensors provided on the moving object.
[0110] As described above, the data generation device 200 in this embodiment generates composite data that shows not only the environment around one moving object, but also the environment including the surroundings of each of multiple moving objects in virtual space, thus enabling a more accurate understanding of a wider range of environments. Furthermore, when each of the multiple moving objects moves, the position of the sensing data of that moving object in virtual space can be changed according to its position in real space after movement. Therefore, even if multiple moving objects move, the composite data can be made to follow their movements.
[0111] (modified version) In the above embodiment, the data generation device 200 is provided on a moving object such as a vehicle 100, but the data generation device 200 may be provided on an external device or server outside the moving object. In this case, the data generation device 200 may generate a wide-area top-view image by acquiring top-view images of each of a plurality of vehicles, including vehicle 100, rather than generating a top-view image of vehicle 100 as in the above embodiment. For example, the wide-area synthesis unit 222 and the image storage unit 221 of the data generation device 200 may be provided on the above-mentioned device or server. Furthermore, the above-mentioned device or server may be at least part of a traffic monitoring cloud.
[0112] Furthermore, in the above embodiment, top-view images are transmitted and received via vehicle-to-vehicle communication, but these top-view images may also be transmitted and received via a traffic monitoring cloud. In this case, the traffic monitoring cloud may request data transmission.
[0113] Furthermore, in the above embodiment, the surrounding vehicle transmits a top-view image to its own vehicle, but the surrounding vehicle may also transmit images obtained by each of the multiple cameras installed on that surrounding vehicle. In other words, instead of a top-view image, the surrounding vehicle may transmit multiple images to its own vehicle that are used to generate the top-view image. In this case, the data generation device 200 of the own vehicle receives the multiple images transmitted from the surrounding vehicle and uses these images to generate a top-view image of the surrounding vehicle. Also, when the surrounding vehicle transmits multiple images obtained by multiple cameras to its own vehicle, it may also transmit the parameter set of each of the multiple cameras to its own vehicle. The parameter set may include parameters indicating the position of the camera on the surrounding vehicle, internal parameters indicating the lens distortion of the camera, and external parameters indicating the orientation of the camera. The data generation device 200 of the own vehicle uses the parameter set to generate a top-view image of the surrounding vehicle.
[0114] Furthermore, surrounding vehicles may transmit only a portion of the top-view image to their own vehicle without transmitting the entire top-view image. For example, the data generation device 200 of the vehicle may specify to the surrounding vehicle the portion of the top-view image of the surrounding vehicle to be transmitted, depending on the relative positional relationship between the vehicle and the surrounding vehicle. At least one of the position, size, and shape of the portion to be transmitted within the entire top-view image may be specified. Alternatively, the surrounding vehicle may identify the portion of its top-view image to be transmitted, depending on its relative positional relationship with the vehicle, and transmit it to the data generation device 200 of the vehicle.
[0115] The position in the above embodiment may be a relative position or an absolute position. The reference point for the relative position may be the position of surrounding vehicles or the vehicle itself. This reference point may differ for each vehicle.
[0116] Furthermore, in the above embodiment, the wide-area top-view image, i.e., the composite data, is displayed by the display unit 230, but it may also be used for signal processing without being displayed. In this case, the data generation device 200 does not need to be equipped with the display unit 230.
[0117] Furthermore, in the above embodiment, the sensing data is a top-view image, but it may be other images, data obtained by a LIDAR (Light Detection and Ranging) or infrared camera, or data obtained by other sensors.
[0118] Furthermore, in the above embodiment, feature points are used in the composite position determination process, but if feature points cannot be detected, the composite position may be determined based on the vehicle's position without using feature points. Also, the vehicle's data generation device 200 acquires position information from surrounding vehicles, but instead of acquiring that position information, it may detect the position of those surrounding vehicles. For example, the data generation device 200 may detect the position of surrounding vehicles using a sensor such as LIDAR or millimeter-wave radar.
[0119] Furthermore, the time managed by each vehicle, including the vehicle itself and surrounding vehicles, may be periodically synchronized using a time synchronization server such as an NTP (Network Time Protocol) server. In the case of platooning, a time synchronization server may be installed in one of the multiple vehicles participating in the platooning. Alternatively, a time synchronization server may be installed for each country or region.
[0120] Furthermore, in the examples of platooning shown in Figures 3 and 4, vehicle C2 is the vehicle itself, and vehicles C1 and C3 are surrounding vehicles. However, vehicle C1 or vehicle C3 may be the vehicle itself, and the other vehicles may be surrounding vehicles. For example, the data generation device 200 of the leading vehicle C1 may present a wide-area top-view image of the entire convoy to the driver of the leading vehicle C1.
[0121] Furthermore, if a pre-set distance between vehicles is used in a platooning system, the position and field of view of each camera placed on each vehicle may be set so that the top-view images of each vehicle overlap, according to the set distance between vehicles.
[0122] Furthermore, the vehicle's data generation device 200 may detect surrounding vehicles traveling around its own vehicle through vehicle-to-vehicle communication or the like, and switch the image presented to the vehicle's driver according to the detection result. For example, the vehicle's data generation device 200 determines whether or not there are surrounding vehicles traveling in front of or behind its own vehicle within a certain distance (for example, within 100m). If the vehicle's data generation device 200 determines that there are surrounding vehicles, it acquires a top-view image of those surrounding vehicles and combines it with the vehicle's top-view image to generate a wide-area top-view image, which is then presented to the vehicle's driver. On the other hand, if the vehicle's data generation device 200 determines that there are no surrounding vehicles, it presents the vehicle's top-view image to the vehicle's driver. In other words, the vehicle's data generation device 200 generates and presents a wide-area top-view image when multiple vehicles, including its own vehicle, are in a predetermined positional relationship, and presents the vehicle's top-view image when the multiple vehicles are not in a predetermined positional relationship.
[0123] Furthermore, the virtual space in the above embodiment is a two-dimensional space such as a top-view image of the vehicle, a lane image, or a lane image in which the top-view image of the vehicle is superimposed, but it may also be a three-dimensional space.
[0124] Furthermore, in the above embodiment, the vehicle is just one example of a moving body, and the moving body may be any object that moves, such as a ship or an aircraft, or any other object besides a vehicle.
[0125] Furthermore, in the above embodiment, the number of sensors provided in the vehicle 100 is four, but this number is not limited to four; it may be three or fewer, or five or more.
[0126] In the above embodiment, each component may be implemented by dedicated hardware or by executing a software program suitable for each component. Here, the software program that implements the data generation device 200, etc., in the above embodiment causes the computer to execute processing according to the flowchart shown in any of Figures 5A, 5B, and 6B.
[0127] Furthermore, each component may be a circuit, as described above. These circuits may form a single circuit as a whole, or they may be separate circuits. Also, each component may be implemented using a general-purpose processor, or it may be implemented using a dedicated processor.
[0128] Furthermore, a process performed by one component may be performed by another component. Also, the order in which processes are executed may be changed, and multiple processes may be executed in parallel.
[0129] The first and second ordinal numbers used in the explanation may be changed as appropriate. Furthermore, ordinal numbers may be newly assigned to or removed from the constituent elements.
[0130] Although the embodiments of the data generation device have been described above based on the above-described embodiments, the embodiments of the data generation device are not limited to those embodiments. Various modifications to the embodiments that a person skilled in the art can conceive of may also be included within the scope of the embodiments of the data generation device, as long as they do not deviate from the spirit of this disclosure.
[0131] (Other embodiments) In the above embodiments, each functional block can typically be implemented by an MPU and memory, etc. Furthermore, the processing performed by each functional block is typically implemented by a program execution unit such as a processor reading and executing software (programs) recorded on a recording medium such as ROM. This software may be distributed by download, etc., or it may be recorded on a recording medium such as semiconductor memory and distributed. Of course, it is also possible to implement each functional block by hardware (dedicated circuitry).
[0132] Furthermore, the processing described in the embodiment may be implemented by centralized processing using a single device (system), or by distributed processing using multiple devices. Also, the processor that executes the above program may be one or multiple. In other words, centralized processing may be performed, or distributed processing may be performed.
[0133] The embodiments of this disclosure are not limited to those described above, and various modifications are possible, which are also included within the scope of the embodiments of this disclosure.
[0134] Furthermore, here we will describe application examples of the video encoding method (image encoding method) or video decoding method (image decoding method) shown in the above embodiment, and a system using the same. The system is characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, and an image encoding and decoding device that includes both. Other configurations in the system can be appropriately modified as needed.
[0135] [Usage example] Figure 7 shows the overall configuration of the content supply system ex100 that realizes the content distribution service. The service area for the communication service is divided into cells of a desired size, and fixed radio stations, base stations ex106, ex107, ex108, ex109, and ex110, are installed in each cell.
[0136] In this content supply system ex100, various devices such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the internet ex101 via an internet service provider ex102 or a communication network ex104, and base stations ex106~ex110. The content supply system ex100 may also connect any combination of the above elements. Each device may be directly or indirectly connected to each other via a telephone network or short-range radio, etc., without going through the base stations ex106~ex110, which are fixed radio stations. In addition, the streaming server ex103 is connected to various devices such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 via the internet ex101, etc. Furthermore, the streaming server ex103 is connected to terminals in a hotspot on an airplane ex117 via satellite ex116.
[0137] Note that instead of base stations ex106~ex110, wireless access points or hotspots may be used. Also, streaming server ex103 may be connected directly to the communication network ex104 without going through the internet ex101 or internet service provider ex102, or it may be connected directly to the airplane ex117 without going through satellite ex116.
[0138] Camera ex113 is a device capable of taking still images and videos, such as a digital camera. Smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handyphone System) that supports mobile communication systems generally known as 2G, 3G, 3.9G, 4G, and the upcoming 5G.
[0139] Home appliance ex118 refers to appliances such as refrigerators or equipment included in household fuel cell cogeneration systems.
[0140] In the content supply system ex100, live streaming becomes possible when a terminal with a shooting function is connected to the streaming server ex103 via a base station ex106 or the like. In live streaming, the terminal (computer ex111, game console ex112, camera ex113, home appliance ex114, smartphone ex115, and terminal inside an airplane ex117, etc.) performs the encoding process described in each of the above embodiments on still images or video content captured by the user using the terminal, multiplexes the video data obtained by encoding with sound data encoded from the sound corresponding to the video, and transmits the obtained data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of this disclosure.
[0141] Meanwhile, the streaming server ex103 streams the content data sent to the requesting client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117, etc., that is capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data. That is, each device functions as an image decoding device according to one aspect of this disclosure.
[0142] [Distributed Processing] Furthermore, the streaming server ex103 may consist of multiple servers or computers that distribute data processing, recording, and distribution. For example, the streaming server ex103 may be implemented using a CDN (Content Delivery Network), where content delivery is achieved through a network connecting numerous edge servers distributed worldwide. In a CDN, the physically closest edge server is dynamically assigned depending on the client. Latency can be reduced by caching and delivering content to the edge server. In addition, if an error occurs or the communication state changes due to an increase in traffic, processing can be distributed among multiple edge servers, the delivery entity can be switched to another edge server, or delivery can be continued by bypassing the failed part of the network, thus enabling high-speed and stable delivery.
[0143] Furthermore, beyond the distributed processing of the distribution itself, the encoding process of the captured data can be performed on each terminal, on the server side, or shared among them. For example, encoding generally involves two processing loops. In the first loop, the complexity or code amount of the image at the frame or scene level is detected. In the second loop, processing is performed to improve encoding efficiency while maintaining image quality. For example, if the terminal performs the first encoding process and the server that receives the content performs the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a request to receive and decode near real time, the first encoded data from the terminal can be received and played back on other terminals, enabling more flexible real-time distribution.
[0144] Another example is the camera ex113, which extracts features from an image, compresses the feature data as metadata, and sends it to the server. The server performs compression according to the meaning of the image, for example, by determining the importance of an object from the features and switching the quantization precision. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during further compression on the server. Alternatively, a simple encoding such as VLC (Variable Length Coding) may be performed on the terminal, and a more computationally intensive encoding such as CABAC (Context-Adaptive Binary Arithmetic Coding) may be performed on the server.
[0145] Another example is a scenario in a stadium, shopping mall, or factory where multiple video data sets of nearly identical scenes may exist, captured by multiple terminals. In such cases, the encoding process is distributed among the multiple terminals that captured the footage, along with other terminals and servers as needed, by assigning encoding tasks to each unit, for example, at the Group of Picture (GOP) level, picture level, or tile level (a division of a picture). This reduces latency and enables more real-time performance.
[0146] Furthermore, since multiple video data sets depict essentially the same scene, the server may manage and / or instruct the video data captured by each terminal to reference each other. Alternatively, the server may receive the encoded data from each terminal, change the reference relationships between the multiple data sets, or correct or replace the pictures themselves and re-encode them. This allows for the creation of a stream with improved quality and efficiency for each individual data set.
[0147] Furthermore, the server may transcode the video data to change its encoding method before distributing it. For example, the server may convert an MPEG-based encoding to a VP-based encoding, or convert H.264 to H.265.
[0148] Thus, the encoding process can be performed by a terminal or one or more servers. Therefore, in the following, the terms "server" or "terminal" will be used to refer to the entity performing the processing, but some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.
[0149] [3D, Multi-angle] In recent years, it has become increasingly common to integrate and utilize images or videos of different scenes, or the same scene, captured from different angles, using multiple cameras ex113 and / or smartphones ex115, which are nearly synchronized with each other. The videos captured by each device are integrated based on the relative positional relationship between the devices, or on areas where feature points contained in the videos coincide, which are acquired separately.
[0150] The server may not only encode 2D video but also encode still images automatically based on scene analysis of the video, or at a time specified by the user, and send them to the receiving terminal. Furthermore, if the server can obtain the relative positional relationship between the shooting terminals, it can generate a 3D shape of the scene based not only on 2D video but also on video of the same scene taken from different angles. The server may also separately encode 3D data generated by a point cloud, or it may select or reconstruct video to send to the receiving terminal from video taken by multiple terminals based on the results of recognizing or tracking a person or object using the 3D data.
[0151] In this way, users can enjoy scenes by arbitrarily selecting each video corresponding to each shooting terminal, or they can enjoy content in which video from an arbitrary viewpoint is extracted from 3D data reconstructed using multiple images or videos. Furthermore, just like the video, sound can also be collected from multiple different angles, and the server may multiplex and transmit sound from a specific angle or space in conjunction with the video.
[0152] In recent years, content that links the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server may create separate viewpoint images for the right and left eyes and perform encoding that allows referencing between the viewpoint images using Multi-View Coding (MVC), or it may encode them as separate streams without referencing each other. When decoding the separate streams, it is advisable to synchronize playback so that the virtual 3D space is reproduced according to the user's viewpoint.
[0153] In the case of AR images, the server superimposes virtual object information from the virtual space onto camera information from the real space, based on its three-dimensional position or the user's viewpoint movement. The decoding device may acquire or store the virtual object information and three-dimensional data, generate a two-dimensional image according to the user's viewpoint movement, and create superimposed data by smoothly stitching them together. Alternatively, the decoding device may send the user's viewpoint movement to the server in addition to requesting virtual object information, and the server may create superimposed data from the three-dimensional data held by the server according to the received viewpoint movement, encode the superimposed data, and distribute it to the decoding device. The superimposed data may have an α value indicating transparency in addition to RGB, and the server may set the α value of parts other than the object created from the three-dimensional data to 0, etc., so that those parts are transparent, and encode the data. Alternatively, the server may set a predetermined RGB value to the background, like chroma keying, and generate data in which parts other than the object are the background color.
[0154] Similarly, the decryption process of the distributed data can be performed on each client terminal, on the server side, or shared between them. For example, one terminal may send a reception request to the server, and other terminals may receive the content corresponding to that request, perform the decryption process, and then transmit the decrypted signal to a device with a display. By distributing the processing and selecting appropriate content regardless of the performance of the communication-capable terminals themselves, it is possible to play back data with good image quality. Another example is that while receiving large image data on a TV or similar device, a portion of the picture, such as tiles, may be decrypted and displayed on the viewer's personal terminal. This allows for sharing the overall picture while allowing users to check their own area of responsibility or areas they want to examine in more detail on their own device.
[0155] In the future, it is expected that content will be seamlessly received by switching appropriate data for the connected communication, using distribution system standards such as MPEG-DASH, in situations where multiple short-range, medium-range, or long-range wireless communications are available both indoors and outdoors. This will allow users to freely select and switch in real time between decoding devices or display devices, such as displays installed indoors or outdoors, as well as their own terminals. Furthermore, decoding can be performed while switching between the decoding terminal and the display terminal based on the user's location information. This will make it possible to display map information on the wall or part of the ground of an adjacent building with a displayable device embedded, while traveling to a destination. It will also be possible to switch the bitrate of the received data based on the ease of access to the encoded data on the network, such as when the encoded data is cached on a server that can be accessed quickly from the receiving terminal, or copied to an edge server in the content delivery service.
[0156] [Scalable encoding] Regarding content switching, we will explain using a scalable stream compressed and encoded using the video encoding method described in each of the embodiments above, as shown in Figure 8. The server may have multiple streams with the same content but different qualities as individual streams, but it may also be configured to switch content by taking advantage of the temporal / spatial scalability of the stream realized by encoding it in layers, as shown in the figure. In other words, the decoding side can freely switch between decoding low-resolution and high-resolution content by deciding which layer to decode according to internal factors such as performance and external factors such as the state of the communication bandwidth. For example, if you want to watch the rest of a video that you were watching on your smartphone ex115 while traveling, on a device such as an internet TV when you get home, that device only needs to decode the same stream to different layers, thus reducing the burden on the server.
[0157] Furthermore, in addition to the configuration described above, in which pictures are encoded for each layer and an enhancement layer exists above the base layer to achieve scalability, the enhancement layer may include metadata based on statistical information of the image, and the decoding side may generate high-quality content by super-resolution the picture in the base layer based on the metadata. Super-resolution may refer to either an improvement in the signal-to-noise ratio at the same resolution or an increase in resolution. The metadata may include information for identifying linear or nonlinear filter coefficients used in the super-resolution process, or information for identifying parameter values in the filtering process, machine learning, or least-squares operation used in the super-resolution process.
[0158] Alternatively, the picture may be divided into tiles or similar structures according to the meaning of objects within the image, and the decoding side may select tiles to decode, thereby decoding only a portion of the area. Furthermore, by storing the attributes of objects (people, cars, balls, etc.) and their positions within the image (coordinate positions within the same image, etc.) as metadata, the decoding side can identify the location of a desired object based on the metadata and determine the tile containing that object. For example, as shown in Figure 9, the metadata is stored using a data storage structure different from pixel data, such as the SEI message in HEVC. This metadata indicates, for example, the position, size, or color of the main object.
[0159] Furthermore, metadata may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. This allows the decryption side to obtain information such as the time when a specific person appears in the video, and by combining this with the picture-level information, it can identify the picture in which the object exists and the object's position within that picture.
[0160] [Web page optimization] Figure 10 shows an example of a web page display screen on a computer ex111, etc. Figure 11 shows an example of a web page display screen on a smartphone ex115, etc. As shown in Figures 10 and 11, a web page may contain multiple linked images, which are links to image content, and their appearance will differ depending on the viewing device. When multiple linked images are visible on the screen, the display device (decoder) will display still images or I-pictures of each content as linked images, display video such as a GIF animation using multiple still images or I-pictures, or receive only the base layer and decode and display the video, until the user explicitly selects a linked image, or until the linked image approaches the center of the screen or the entire linked image is within the screen.
[0161] When a linked image is selected by the user, the display device prioritizes decoding the base layer. If the HTML of the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. Furthermore, to ensure real-time performance, before selection or when bandwidth is very limited, the display device can decode and display only forward-referenced pictures (I-pictures, P-pictures, and B-pictures that only use forward references), thereby reducing the delay between the decoding time and display time of the first picture (the delay from the start of content decoding to the start of display). Alternatively, the display device may deliberately ignore the reference relationships between pictures and roughly decode all B-pictures and P-pictures using forward references, then perform normal decoding as time passes and more pictures are received.
[0162] [Autonomous driving] Furthermore, when transmitting and receiving still images or video data such as 2D or 3D map information for autonomous driving or driving assistance of a vehicle, the receiving terminal may receive metadata such as weather or construction information in addition to image data belonging to one or more layers, and decode these in association with each other. The metadata may belong to a layer, or it may simply be multiplexed with the image data.
[0163] In this case, since the vehicle, drone, or airplane containing the receiving terminal is in motion, the receiving terminal can transmit its location information when a reception request is made, enabling seamless reception and decoding while switching between base stations ex106 to ex110. Furthermore, the receiving terminal can dynamically switch how much metadata is received or how much map information is updated, depending on the user's selection, the user's situation, or the state of the communication bandwidth.
[0164] As described above, the content supply system ex100 allows the client to receive, decode, and play back encoded information transmitted by the user in real time.
[0165] [Distribution of personal content] Furthermore, the ex100 content delivery system allows for unicast or multicast distribution of not only high-definition, long-duration content from video distribution companies, but also low-definition, short-duration content from individuals. It is also expected that the amount of such individual content will continue to increase. To improve the quality of individual content, the server may perform editing before encoding. This can be achieved, for example, with the following configuration.
[0166] During shooting, or after shooting, the server performs recognition processing such as detecting shooting errors, searching for scenes, analyzing semantics, and detecting objects from the original images or encoded data in real time. Based on the recognition results, the server manually or automatically edits the images, correcting out-of-focus or shaky images, deleting less important scenes such as those with lower brightness or out of focus compared to other pictures, emphasizing object edges, and changing color tones. The server then encodes the edited data based on the editing results. It is also known that viewership decreases if the shooting time is too long, so the server may automatically clip scenes with little movement, as well as less important scenes, based on the image processing results, to ensure that the content falls within a specific time range according to the shooting time. Alternatively, the server may generate and encode a digest based on the results of the semantic analysis of the scenes.
[0167] Furthermore, personal content may contain elements that infringe on copyright, moral rights, or portrait rights, and the scope of sharing may exceed the intended scope, which can be inconvenient for the individual. Therefore, for example, the server may intentionally change the image to one that is out of focus, such as the faces of people at the edges of the screen or the interior of a house, before encoding. The server may also recognize whether the face of a person other than those previously registered is visible in the image to be encoded, and if so, it may apply a mosaic effect to the face. Alternatively, as a pre- or post-processing step before encoding, the user can specify a person or background area that they want to process from a copyright perspective, and the server can replace the specified area with a different image or blur the focus. In the case of a person, the server can track the person in a video and replace the image of their face.
[0168] Furthermore, because viewing personal content with small data volumes requires real-time processing, depending on the bandwidth, the decoder prioritizes receiving, decoding, and playing the base layer first. During this time, the decoder can receive the enhancement layer, and if playback is looped or if the content is played more than once, it may play the high-quality video including the enhancement layer. With a stream that uses this scalable encoding, it is possible to provide an experience where the video is rough when unselected or at the beginning of viewing, but gradually the stream becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if the rough stream played the first time and the second stream encoded by referencing the first video are configured as a single stream.
[0169] [Other usage examples] Furthermore, these encoding or decoding processes are generally performed by the LSIex500 present in each terminal. The LSIex500 may be a single chip or a multi-chip configuration. Alternatively, video encoding or decoding software may be embedded in some recording medium (such as a CD-ROM, flexible disk, or hard disk) that can be read by a computer ex111, and the encoding or decoding process may be performed using that software. In addition, if the smartphone ex115 has a camera, video data acquired by that camera may be transmitted. In this case, the video data is data encoded by the LSIex500 present in the smartphone ex115.
[0170] The LSIex500 may also be configured to be activated by downloading application software. In this case, the terminal first determines whether it supports the content encoding method or whether it has the capability to perform the specific service. If the terminal does not support the content encoding method or does not have the capability to perform the specific service, the terminal downloads the codec or application software, and then acquires and plays the content.
[0171] Furthermore, not only the content supply system ex100 via the Internet ex101, but also digital broadcasting systems can incorporate at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of each of the above embodiments. While the content supply system ex100 has a configuration that is more suited to multicast than unicast, as it transmits and receives multiplexed data with video and sound multiplexed onto broadcast radio waves using satellites, etc., the encoding and decoding processes are similar and can be applied in the same way.
[0172] [Hardware configuration] Figure 12 shows the smartphone ex115. Figure 13 shows an example of the configuration of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with the base station ex110, a camera unit ex465 capable of taking video and still images, and a display unit ex458 that displays video captured by the camera unit ex465 and data decoded from video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466, such as a touch panel, an audio output unit ex457, such as a speaker for outputting voice or sound, an audio input unit ex456, such as a microphone for inputting voice, a memory unit ex467 capable of storing captured video or still images, recorded audio, received video or still images, encoded data such as emails, or decoded data, and a slot unit ex464, which is an interface unit with SIM ex468 for identifying the user and authenticating access to various data, including the network. External memory may be used instead of the memory unit ex467.
[0173] Furthermore, the main control unit ex460, which comprehensively controls the display unit ex458 and the operation unit ex466, is connected via the bus ex470 to the power supply circuit unit ex461, the operation input control unit ex462, the video signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / decompression unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467.
[0174] The power supply circuit unit ex461, when the power key is turned on by the user, supplies power from the battery pack to each component, thereby starting up the smartphone ex115 and making it operational.
[0175] The smartphone ex115 performs tasks such as phone calls and data communication based on the control of the main control unit ex460, which has a CPU, ROM, RAM, etc. During a call, the audio signal picked up by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454, which is then subjected to spread spectrum processing by the modulation / demodulation unit ex452, and after digital-to-analog conversion and frequency conversion processing by the transmission / reception unit ex451, it is transmitted via the antenna ex450. Similarly, received data is amplified, subjected to frequency conversion and analog-to-digital conversion processing, despread spectrum processing by the modulation / demodulation unit ex452, converted into an analog audio signal by the audio signal processing unit ex454, and then output from the audio output unit ex457. In data communication mode, text, still images, or video data are sent to the main control unit ex460 via the operation input control unit ex462 by the operation unit ex466 of the main unit, and transmission and reception processing is performed in the same manner. When transmitting video, still images, or video and audio in data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the video encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / decoding unit ex453. The audio signal processing unit ex454 encodes the audio signal picked up by the audio input unit ex456 while the camera unit ex465 is capturing video or still images, and sends the encoded audio data to the multiplexing / decoding unit ex453. The multiplexing / decoding unit ex453 multiplexes the encoded video data and encoded audio data in a predetermined manner, performs modulation and conversion processing in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits the data via the antenna ex450.
[0176] When receiving video attached to an email or chat, or video linked to a webpage, etc., the multiplexing / decomposition unit ex453 separates the multiplexed data received via antenna ex450 to decode the multiplexed data, dividing it into a video data bitstream and an audio data bitstream. It then supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronization bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method shown in each embodiment above, and displays the video or still image contained in the linked video file from the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. However, since real-time streaming is widespread, there may be situations where audio playback is socially inappropriate depending on the user's circumstances. Therefore, as an initial setting, it is preferable to have a configuration that plays only video data and not audio signals. Audio may be synchronized and played only when the user performs an action, such as clicking on video data.
[0177] Furthermore, although the smartphone ex115 was used as an example here, there are three possible implementation formats for terminals: a transceiver-type terminal that has both an encoder and a decoder, a transmitting terminal that has only an encoder, and a receiving terminal that has only a decoder. In addition, although it was explained that multiplexed data, in which audio data etc. is multiplexed with video data, is received or transmitted in a digital broadcasting system, the multiplexed data may also include text data related to the video in addition to audio data, or the video data itself may be received or transmitted instead of multiplexed data.
[0178] Although it was explained that the main control unit ex460, including the CPU, controls the encoding or decoding process, terminals often also have a GPU. Therefore, a configuration that leverages the GPU's performance to process a wide area at once using memory shared by the CPU and GPU, or memory whose addresses are managed so that it can be used in common, is also possible. This can shorten the encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient to perform motion detection, deblocking filters, SAO (Sample Adaptive Offset), and transformation / quantization processes at once on the GPU, rather than on the CPU, in units such as pictures. [Industrial applicability]
[0179] The data generation device disclosed herein has the potential for further improvement and, for example, can be installed in a vehicle and used as an in-vehicle device to assist in the driving of that vehicle, thus having high utility value. [Explanation of Symbols]
[0180] 100 vehicles 101 First Sensor 102 Second Sensor 103 Third Sensor 104 Fourth Sensor 110 Position and direction detection unit 200 Data Generation Devices 201 Circuit 202 memory 210 Acquisition Department 211 First Communications Department 212 Receiving Control Unit 213 Data receiving unit 214 First Format Conversion Unit 220 Synthesis section 221 Image storage unit 222 Wide-area synthesis section 223 Top View Image Generation Unit 230 Display section 240 Output section 241 Second Communications Department 242 Transmission Control Unit 243 Data transmission section 244 Second Format Conversion Section
Claims
1. Circuits and, The circuit comprises a memory connected to the aforementioned circuit, In operation, the aforementioned circuit A first image is obtained using a camera mounted on the first mobile device. A second image is obtained using a camera mounted on the second mobile device. A composite image is generated from the first image and the second image. In generating the aforementioned composite image, Based on the positional information of the first and second moving objects and the matching result between the feature points extracted from the first image and the feature points extracted from the second image, the composite position of the first and second images is determined. The first image and the second image are, respectively, a portion of the top view image of the first moving object and a portion of the top view image of the second moving object. The aforementioned top-view image is an image of the surroundings including the moving object corresponding to the top-view image, viewed from above the moving object. The portion of the top-view image is a part that is identified according to the relative positional relationship between the first moving object and the second moving object. Image synthesis device.
2. The aforementioned circuit is The orientation of the first image and the second image is determined, In determining the composite position of the first and second images, Based on the relative positions of the first moving body and the second moving body, candidate composite positions of the first and second images are determined, and based on the matching results, the determined candidate composite positions are refined to determine the composite position of the first and second images. In generating the aforementioned composite image, The orientation of the first image and the second image is changed to the determined orientation, and the first image and the second image are superimposed on the determined composite position. The composite image generation apparatus according to claim 1.
3. The aforementioned circuit is Determine whether the first moving body and the second moving body are in a predetermined positional relationship. (a) If it is determined that the predetermined positional relationship is as described above, The acquisition of the first image and the second image, and the generation of the composite image are performed. (b) If it is determined that the positional relationship is not as predetermined as described above, Without generating the aforementioned composite image, the first image or the second image is acquired and displayed. The composite image generation apparatus according to claim 1.
4. The position information indicates the positions of the first and second moving objects at the time the first and second images were taken. The composite image generation apparatus according to claim 1.
5. The circuit further, In generating the aforementioned composite image, Based on the direction of travel information indicating the direction of travel of the first moving body and the second moving body, the composite position of the first image and the second image is determined. The composite image generation apparatus according to claim 1.
6. The first moving body and the second moving body are vehicles. The composite image generation apparatus according to claim 1.
7. The first image and the second image include regions that overlap with each other. The composite image generation apparatus according to claim 6.
8. A first image is obtained using a camera mounted on the first mobile device. A second image is obtained using a camera mounted on the second mobile device. A composite image is generated from the first image and the second image. In generating the aforementioned composite image, Based on the positional information of the first and second moving objects and the matching result between the feature points extracted from the first image and the feature points extracted from the second image, the composite position of the first and second images is determined. The first image and the second image are, respectively, a portion of the top view image of the first moving object and a portion of the top view image of the second moving object. The aforementioned top-view image is an image of the surroundings including the moving object corresponding to the top-view image, viewed from above the moving object. The portion of the top-view image is a part that is identified according to the relative positional relationship between the first moving object and the second moving object. A method for generating composite images.
Citation Information
Patent Citations
Image generation apparatus and image generation program
JP2014089513A