Communication systems, communication devices, and communication methods

The communication system addresses bandwidth limitations by dividing high-resolution videos into importance-based elemental stages with adjusted density, reducing transmission needs while maintaining quality.

JP2026077369AActive Publication Date: 2026-05-13LENOVO (SINGAPORE) PTE LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LENOVO (SINGAPORE) PTE LTD
Filing Date
2024-10-25
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Transmitting high-resolution videos, such as 4K videos, requires a large amount of communication capacity, and existing methods like ITU-T 264.H may not suffice under limited bandwidth conditions.

Method used

A communication system that divides the video into multiple stages of elemental videos with varying importance levels, adjusting their information density, and transmits these stages to another device for reconstruction, where the higher importance stages have higher density, using a first device and a second device with video processing units to synthesize the restored video.

Benefits of technology

Reduces the required transmission capacity without hindering communication quality by optimizing the information density of different video stages, allowing efficient transmission of high-resolution videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026077369000001_ABST
    Figure 2026077369000001_ABST
Patent Text Reader

Abstract

Reduce the overall transmission capacity required by the system to avoid hindering communication. [Solution] The system comprises at least a first device and a second device. The first device divides the original video into multiple stages of elemental video with different levels of importance based on the appearance of the subject, adjusts the information density of the elemental video so that the level of importance increases relatively, and transmits the multiple stages of elemental video to the second device. The second device receives the multiple stages of elemental video from the first device and synthesizes the multiple stages of elemental video to reconstruct a restored video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a communication system, a communication device, and a communication method, for example, a video communication system for having a conversation among multiple people.

Background Art

[0002] In recent years, video communication systems that share videos acquired at multiple locations have become widespread. Also, due to the development of imaging equipment and communication technologies, some video communication systems can implement high-definition video conferences that can handle high-resolution videos. It is expected that non-verbal communication will be facilitated by high-resolution videos.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, transmitting high-resolution videos such as 4K videos requires a large amount of communication capacity (in the case of wireless communication, bandwidth). Even when high-resolution videos are compression-encoded using a video encoding method such as ITU-T 264.H, they may not be transmitted as they are under limited capacity. For example, the information processing device described in Patent Document 1 acquires encoded video data captured by a camera from the camera via a network, performs decoding processing on the acquired video data, and outputs the decoded processed video data to a video processing application via a driver. When the amount of stored video data is equal to or greater than a predetermined threshold, or when the difference between the time information included in the latest video data and the time information included in the oldest video data among the stored video data is equal to or greater than a predetermined threshold, a warning is displayed. [Means for solving the problem]

[0005] This invention was made to solve the above problems, and a communication system according to one embodiment is a communication system comprising at least a first device and a second device, wherein the first device comprises a first video processing unit that divides a source video into multiple stages of elemental video of different importance based on the form of the subject, and adjusts the information density of the elemental video so that the higher the importance of the stage, the higher the relative information density of the elemental video, and a first communication processing unit that transmits the multiple stages of elemental video to the second device, and the second device comprises a second communication processing unit that receives the multiple stages of elemental video from the first device, and a second video processing unit that synthesizes the multiple stages of elemental video to reconstruct a restored video.

[0006] In the above-described communication system, the first video processing unit may detect feature information indicating the morphological features and position of the subject from the original video, and the second video processing unit may reconstruct the restored video by superimposing multiple stages of the elemental video so that the morphological features are positioned at the locations indicated by the feature information.

[0007] In the above-described communication system, the second video processing unit may convert the multiple-stage elemental video so that its information density is equal to the highest information density among the multiple-stage elemental video, and then reconstruct the restored video by superimposing the converted multiple-stage elemental video.

[0008] In the above-described communication system, the second video processing unit may reconstruct the restored video by superimposing multiple stages of the converted video elements, such that the video elements at stages of higher importance are given priority.

[0009] In the above-described communication system, the first video processing unit may detect the amount of movement of a subject in a specific manner from the original video, and if the amount of movement is outside a predetermined range of movement, it may determine the element video representing the subject as an element video of the lowest level, which is the least important level. If the amount of movement is within the range of movement, it may determine the element video representing the subject as a candidate for an element video of a higher level, which is more important than the element video of the lowest level.

[0010] In the above-described communication system, the first video processing unit may detect the size of an image of a subject of a specific nature from the original video, and if the size is outside a predetermined range of size, it may determine the element video representing the subject as a basic video of the lowest level, which is the least important level. If the size is within the range of size, it may determine the element video representing the subject as a candidate for a higher-level element video, which is more important than the lowest-level element video.

[0011] In the above-described communication system, the subject of the particular form may be the upper body of a person or a display medium that displays content.

[0012] The communication device according to the second embodiment comprises an image processing unit that divides the original image into a first set of multiple elemental images of different importance based on the appearance of the subject, and adjusts the information density of the elemental images so that the density increases with increasing importance, and a communication processing unit that transmits the first set of multiple elemental images to another device, wherein the communication processing unit receives at least a second set of multiple elemental images from the other device, and the image processing unit synthesizes the second set of multiple elemental images to reconstruct a restored image.

[0013] A communication method according to the third embodiment is a communication system comprising at least a first device and a second device, wherein the first device performs the steps of: dividing the original image into multiple stages of elemental images of different importance based on the appearance of the subject; adjusting the information density of the elemental images so that the higher the importance of the stage, the higher the relative information density; and transmitting the multiple stages of elemental images to the second device; and the second device performs the steps of: receiving the multiple stages of elemental images from the first device; and synthesizing the multiple stages of elemental images to reconstruct a restored image. [Effects of the Invention]

[0014] According to the embodiment of the present invention, the transmission capacity required for the entire system can be reduced without hindering communication. [Brief explanation of the drawing]

[0015] [Figure 1] This is a schematic block diagram showing an example configuration of the communication system according to this embodiment. [Figure 2] This is a schematic block diagram showing an example of the hardware configuration of the terminal device according to this embodiment. [Figure 3] This is a schematic block diagram showing an example of the functional configuration of the terminal device according to this embodiment. [Figure 4] This is a flowchart illustrating the video communication processing according to this embodiment. [Figure 5] This figure illustrates a still image extracted from the original image. [Figure 6] This diagram illustrates elemental images extracted from a still image. [Figure 7] This figure shows an example of feature information detection from a still image. [Figure 8] This diagram illustrates the basic still images that make up the basic video. [Figure 9] This figure illustrates a reconstructed still image created by combining an elemental image with a basic still image. [Figure 10] This is an explanatory diagram illustrating the compositing of elemental video and base video. [Figure 11]This is a flowchart exemplifying the main part determination process according to this embodiment.

Embodiment for Carrying Out the Invention

[0016] Hereinafter, embodiments of the present application will be described with reference to the drawings. A configuration example of the communication system S1 according to this embodiment will be described. FIG. 1 is a schematic block diagram showing a configuration example of the communication system S1 according to this embodiment. The communication system S1 is configured to include a plurality of terminal devices 1. The plurality of terminal devices 1 are connected to each other via a network NW so as to be able to transmit and receive various data. The network NW may be any one or a combination of a plurality of the Internet, a public communication network, an in-house communication network (LAN: Local Area Network), a virtual private network (VPN: Virtual Personal Network), a dedicated line, etc.

[0017] In the example of FIG. 1, two terminal devices 1-1 and 1-2 are shown. The two terminal devices 1-1 and 1-2 are distinguished by attaching sub-numbers -1 or -2. In the present application, the sub-numbers may be omitted. The number of terminal devices 1 related to one communication may also be called the number of connections, the number of sites, etc. The number of connections is not limited to two, and may be an unspecified number of three or more, or may be fixed to a specific number.

[0018] The communication system S1 is configured as, for example, a conference system, a video call system, etc. The data transmitted and received between the plurality of terminal devices 1 includes video data. Communication of intentions is achieved among the users of the individual terminal devices 1 using the data transmitted and received. A dedicated communication server is connected to the network NW, and data may be transmitted between the plurality of terminal devices 1 via the communication server.

[0019] At least one of the multiple terminal devices 1 (for example, terminal device 1-1) acquires the video to be transmitted (referred to as the "original video" in this application). In this application, the terminal device 1 that transmits the video may be referred to as the "first device". The first device divides the original video into multiple stages of elemental video with different levels of importance based on the appearance of the subject, and adjusts the information density of the elemental video so that the level of importance increases relatively. The first device transmits the adjusted multiple stages of elemental video to other terminal devices (for example, terminal device 1-2). In this application, the terminal device 1 that receives the video may be referred to as the "second device". The second device receives multiple stages of elemental video from the first device, synthesizes the received multiple stages of elemental video, and reconstructs a restored image. The second device then presents the reconstructed restored image.

[0020] Here, the first device refers to terminal device 1 that transmits video, and the second device refers to terminal device 1 that receives video, and does not refer to the hardware configuration or other functional configuration of individual terminal devices 1. Each terminal device 1 may (1) perform the functions of the first device and not the functions of the second device, (2) perform the functions of the second device and not the first device, or (3) perform the functions of both the first and second devices. In the example of the functional configuration of terminal device 1 described later, it has the functions of both the first and second devices.

[0021] Next, an example of the configuration of the terminal device 1 according to this embodiment will be described. Figure 2 is a schematic block diagram showing an example of the hardware configuration of the terminal device 1 according to this embodiment. The terminal device 1 may be a general-purpose information terminal device such as a personal computer (PC), a tablet terminal device, or a multi-function mobile phone (including a so-called smartphone), or it may be a terminal device dedicated to conferences or calls. In the following description, the case where the terminal device 1 is a PC will be used as an example.

[0022] The terminal device 1 includes a host system 10, a ROM (Read Only Memory) 22, an auxiliary storage device 23, a display 24, a camera 25, an audio system 26, a communication module 27, an input / output interface 28, an embedded controller (EC) 31, an input device 32, a power supply circuit 33, and a power switch 36.

[0023] The host system 10 is the core computer system of the terminal device 1. The host system 10 comprises a processor, main memory, and a chipset. In this application, the hardware constituting the host system 10 may be referred to as the "host device." The processor is the core processing unit that controls the operation of the entire terminal device 1. The processor is, for example, a CPU (Central Processing Unit). The processor is the core processing unit that executes arithmetic processing instructed by various commands written in the software (program). In this application, the execution of processing instructed by commands written in the program may be referred to as "executing the program" or "program execution."

[0024] The processor may include a CPU as well as a GPU. The GPU is a processing unit primarily used to implement functions related to image display. The GPU processes drawing instructions issued by the CPU (image processing) and outputs display data showing the obtained display information to the display 24. The GPU may be integrated with the CPU and formed on the same core, or it may be formed on a separate core from the CPU.

[0025] Main memory is writable memory used as a reading area for the processor's executable program or as a working area for writing processing data for the executable program. Main memory is composed of, for example, multiple DRAM (Dynamic Random Access Memory) chips. The processor and main memory constitute the minimum hardware required for the host system 10.

[0026] The chipset includes multiple controllers and is capable of connecting to multiple devices to input and output various types of data. The controllers on the chipset 21 may be, for example, USB (Universal Serial Bus), SPI (Serial Peripheral Interface) bus, PCI-Express bus, etc.

[0027] ROM22 primarily stores firmware. The firmware stored in ROM22 includes the BIOS (Basic Input-Output System) and other firmware specific to individual devices. ROM22 is composed of rewritable non-volatile memory such as EEPROM (Electrically Erasable Programmable Read Only Memory) and flash ROM.

[0028] The auxiliary storage device 23 stores various data used in the processing of the host system 10, various data acquired through such processing, or various programs. The auxiliary storage device 23 may be, for example, an SSD (Solid State Drive) or an HDD (Hard-disk Drive).

[0029] The display 24 displays a screen based on display data input from the host system 10. The display 24 may be, for example, a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display.

[0030] Camera 25, following the control of the host system 10, captures images representing subjects appearing within its field of view. Camera 25 is a video camera capable of capturing video, i.e., moving images. Video generally consists of still images taken at regular intervals in chronological order. Camera 25 outputs video data showing the captured video to the host system 10.

[0031] The audio system 26 includes an audio codec and performs input and output of audio data. The audio codec, under the control of the host system 10, converts the analog input audio signal received from the microphone into digital input audio data and outputs the resulting input audio data to the host system 10. The microphone detects the sound propagating within itself and outputs an input audio signal indicating the detected sound to the audio system 26. The audio codec includes a decoder that converts the digital output audio data output from the host system 10 into an analog output audio signal and outputs the resulting output audio signal to the speaker. The speaker presents sound based on the output audio signal received from the audio system 26. The microphone and speaker may be provided on the terminal device 1, one or both may be detachably connected to the terminal device 1, or they may be separate from the terminal device 1.

[0032] The communication module 27 connects to a communication network, enabling it to send and receive various types of data wirelessly or via a wired connection. The communication module 27 communicates various types of data with other devices connected to the communication network. The communication module 27 is, for example, a wireless LAN module that connects to a wireless LAN.

[0033] The I / F28 connects to various devices for data input and output via wired or wireless connections. For example, the I / F28 includes a USB connector for wired data input and output in accordance with USB specifications.

[0034] EC31 is a controller that monitors and controls the operation of various devices connected to it, regardless of the operating status of the host system 10. EC31 has a separate CPU, ROM, RAM, timer, and input / output interface from the host system 10. Devices with slower data transfer speeds than those connected to the host system 10 can be connected to EC31. In the example shown in Figure 2, an input device 32, a power supply circuit 33, and a power switch 36 are connected to EC31.

[0035] The EC31 reads a predetermined firmware from its own ROM, executes the read firmware, and provides its function. Alternatively, the firmware may be pre-stored in ROM22 instead of its own ROM, and the read firmware may be executed. However, ROM22 must also be started when the EC31 starts up.

[0036] The input device 32 detects user operations, generates an operation signal according to the detected operation, and outputs it to EC31. The input device 32 may be, for example, a keyboard, a touchpad, or any other.

[0037] The power supply circuit 33 includes a voltage converter and a charger. The voltage converter converts the voltage of DC power supplied from an external power source or battery (not shown) to the voltage required for the operation of each device constituting the terminal device 1, and supplies power with the converted voltage to the receiving device. The power supply circuit 33 performs power supply to the device according to the control of EC31. The charger charges the battery with any remaining power from the external power source that is not consumed by each device. If no power is supplied from the external power source, or if the power supplied from the external power source does not meet the demand, the charger supplies power discharged from the battery to each device. The battery charges with power supplied from the power supply circuit 33, or discharges power stored in itself to the power supply circuit 33. The battery may be, for example, a lithium-ion battery, a sodium-ion battery, or any other type.

[0038] Each time a press operation is received, the power switch 36 controls the power supply state to the host system 10 to either power ON or power OFF. When a press operation is received, the power switch 36 outputs a press signal to EC31 indicating that a press has been made. When the terminal device 1 is powered off and a press signal is input from the power switch 36, EC31 instructs the power supply circuit 33 to start supplying power to each device of the terminal device 1 (power on). When power is supplied to the terminal device 1 and a press signal is input from the power switch 36, EC31 causes the host system 10 to perform a shutdown process.

[0039] Next, an example of the functional configuration of terminal device 1 will be described. The functions of terminal device 1 are realized by the host system 10's processor executing various programs in cooperation with the main memory and other hardware. Figure 3 is a schematic block diagram showing an example of the functional configuration of terminal device 1 according to this embodiment. Terminal device 1 includes a video processing unit 12, a communication processing unit 14, and an output processing unit 16. The video processing unit 12 includes an analysis unit 122, a multiplexing unit 124, and a synthesis unit 126.

[0040] The analysis unit 122 receives video data from the camera 25 (Figure 2). The analysis unit 122 analyzes the original video shown in the video data to detect the display area of ​​a subject of a specific type. The type of subject, or a combination of that type and its operating state, is determined. For example, the analysis unit 122 performs known image recognition processing on the original video, which is composed of still images of multiple frames, to determine the type of subject appearing in the original video and the time change of its display area. In the image recognition processing, machine learning models such as LSTM (Long Short-Term Memory) and RNN (Recurrent Neural Network) are used. Based on the determined subject type, the analysis unit 122 divides the image into a predetermined number of stages. The analysis unit 122 has a corresponding stage set in advance for each subject type and classifies each subject type into one of the multiple stages. The importance of each stage in communication between users differs.

[0041] For example, a person's face is more important than other parts of the person or other subjects, regardless of whether it is moving or not. Within a person's face, certain organs, particularly the eyes and lips, are more important than other organs. Even among other organs, multiple parts of the head in motion are more important than when stationary. Changes in the relative positions of these multiple parts represent the person's facial expression. Furthermore, a person's head, hands, or arms in motion are more important than when stationary. Movements of the head, hands, or arms form gestures. In other words, the posture of a person's upper body provides clues for nonverbal communication.

[0042] Furthermore, display media that allow content to be displayed on a surface through user operation are of greater importance than other subjects. This is because such display media can be used as a means to complement information transmission in video communication between users. The displayed content consists of one or a combination of characters, symbols, or figures. Examples of such display media include touch panels and bulletin boards. A touch panel may be a standalone device or part of an information terminal device that primarily performs other functions, such as a smartphone. A bulletin board may be in any form, such as a blackboard, whiteboard, or notebook, and its display mechanism may be electronic or non-electronic. Such display media may, for example, have a display surface that shows the trajectory of a writing motion (drawing input).

[0043] However, display media in which content is not displayed through operation are of low importance. More specifically, display media that are held by a person or whose content display surface is indicated may be judged as more important than display media that are not held or whose display surface is not indicated. Furthermore, objects that form the background, such as walls and windows, equipment other than the display medium, and clothing may be classified as the least important. This is because these objects do not contribute to communication between users. If there are multiple subjects in the least important stage in the original video, the analysis unit 122 may not distinguish between these multiple subjects and may detect the collection of these multiple subjects as the background.

[0044] The analysis unit 122 analyzes the original video shown in the video data and detects feature information indicating predetermined morphological features appearing on the subject. The analysis unit 122 may explicitly include information on the location of the detected morphological features in the feature information. The analysis unit 122 detects feature points and / or edges as feature information using known image processing techniques. Feature information detection may be performed as part of the image recognition process using a machine learning model, or it may be performed independently. The analysis unit 122 may identify element images such that each subject contains one or more feature points. The analysis unit 122 may identify element images such that the outer edge of each subject contains part or all of an edge.

[0045] Feature points are detected as minute areas that differ in color or shade from their surroundings. Feature points are sometimes called landmarks. A feature point corresponds to a group of regions where a predetermined number of pixels are adjacent spatially, and the diameter is less than or equal to a predetermined size, where the gradient of the signal value of each pixel is steeper than a predetermined reference value. The signal value to be processed may be either a luminance value or a color signal value. Typical feature points on the face include dimples, pockmarks, moles, crow's feet, corners of the mouth, and cleft lip.

[0046] Edges are defined as contours or lines that exhibit significant changes in color or shading. An edge is a region where a predetermined number of pixels are spatially adjacent, where the gradient of the signal value for each pixel is steeper than a predetermined reference value, and where the length in a specific direction is longer than a predetermined length, and the width in the direction intersecting that length is sufficiently narrower than the length. Typical examples of edges include the jawline forming the base of the head, the lateral edges, the crown or outer edge of the hair, and the outer edge of the hand or arm.

[0047] The analysis unit 122 adjusts the information density of each elemental image in multiple stages using known image processing techniques so that the information density of the elemental image becomes relatively higher for stages of higher importance. The information density of an image is determined by its resolution, bit depth, and frame rate. Generally, resolution is indicated by the number of pixels per frame. The more pixels per frame, the higher the information density. Bit depth is the number of bits used to represent the signal value for each pixel. The higher the bit depth, the higher the information density. Frame rate corresponds to the number of frames per second. The higher the frame rate, the shorter the frame interval (period) becomes, and the higher the information density. However, the analysis unit 122 reduces the total information amount of the multiple elemental images to less than the information amount of the original image. Therefore, the rate of reduction in information density is higher for stages of lower importance. The analysis unit 122 outputs multiple stages of elemental video and feature information to the multiplexing unit 124.

[0048] The multiplexing unit 124 multiplexes the individual elemental images and feature information input from the analysis unit 122. The multiplexing unit 124 generates encoded data by performing encoding processing on each of the multiple stages of elemental video using a predetermined video encoding scheme. As the predetermined video encoding scheme, the multiplexing unit 124 can use, for example, the scheme specified in ITU-T H.264 (AVC: Advanced Video Coding) or the scheme specified in ITU-T H.265 (VVC: Versatile Video Coding). The multiplexing unit 124 generates multiplexed data by associating the encoded data of each of the multiple stages with feature information and multiplexing them together. The multiplexing unit 124 outputs the generated multiplexed data to the communication processing unit 14. The output multiplexed data is transmitted to the terminal device 1 (sometimes referred to as the "recipient device" in this application) using the communication processing unit 14.

[0049] The synthesis unit 126 receives multiplexed data received from the other device via the communication processing unit 14. The synthesis unit 126 separates the multiplexed data and feature information of each of the multiple stages from the input multiplexed data. The synthesis unit 126 then performs a decoding process on the separated encoded data of each stage using a predetermined video decoding method to reconstruct the elemental video. The predetermined video decoding method can be any method that corresponds to the video encoding method used to convert the data to encoded data.

[0050] The synthesis unit 126 identifies the positions such that the morphological features of the subject corresponding to the positions indicated in the feature information are arranged, and then reconstructs the restored image by superimposing multiple stages of elemental images. Before superimposing the multiple stages of elemental images, the synthesis unit 126 uses known image processing techniques to convert the information density of the multiple stages of elemental images to a single common format, and acquires the converted elemental image as the converted elemental image. The synthesis unit 126 spatially and temporally interpolates the signal values ​​of each pixel of the elemental images in each frame constituting the first stage of elemental images so that the resolution, bit depth, and frame rate per unit area of ​​a certain first stage of elemental image are equal to the resolution, bit depth, and frame rate of the more important second stage of elemental image. As the single information density after commonization, for example, the highest information density among the multiple stages of information density may be applied. By making the information density a single format across stages, it becomes easier to manipulate the pixel values ​​for each frame and each pixel. It is also convenient for identifying the positions of elemental images based on morphological features.

[0051] A region of a first-stage transformation element image whose position has been identified may overlap with a portion of a second-stage transformation element image, which is another stage. Therefore, when the compositing unit 126 superimposes multiple stages of transformation element images whose positions have been identified, it prioritizes the transformation element images of stages with higher importance. That is, if signal values ​​are set for each of two or more stages of transformation element images for a given pixel, the compositing unit 126 adopts the signal value related to the transformation element image of the stage with the highest importance and rejects the signal values ​​related to the transformation element images of the other stages. In this case, the compositing unit 126 can simply overwrite the lowest-stage transformation element image, which is the least important stage, with the transformation element images of stages with higher importance.

[0052] It should be noted that the video encoding and decoding methods exemplified above can encode and decode videos with rectangular regions, but cannot process videos with arbitrary shapes. Therefore, the processing target is limited to videos with a rectangular region inscribed within the subject. As a result, in the restored video, the gap between the rectangular region related to the superimposed element video and the subject may be visible, which may cause discomfort to the user.

[0053] Therefore, the synthesis unit 126 may perform image recognition processing on a rectangular region containing element images of higher importance than the lowest level to extract element images representing the subject and eliminate the gaps between them and the rectangular region. Using the above method, the synthesis unit 126 superimposes the extracted element images to synthesize the reconstructed image. The gaps between the rectangular region and the element images will no longer appear in the reconstructed image. The synthesis unit 126 outputs video data showing the synthesized restored image to the output processing unit 16.

[0054] The communication processing unit 14 identifies the terminal device 1, which is connected to the network NW and will be the remote device. The remote device may be identified using any of the following identification information: URL (Universal Resource Locator), SIP-URI (Session Initiation Protocol-Universal Resource Identifier), etc. For example, when a connection request is input to the local device from a remote device, the communication processing unit 14 displays an inquiry screen on the display 24 to ask whether or not to allow the connection with the remote device. The communication processing unit 14 determines that the connection is allowed when the allow button on the inquiry screen is indicated by an operation signal input from the input device 32. The communication processing unit 14 then sends an acknowledgment to the remote device that sent the connection request. At this point, a connection is established between the local device and the remote device.

[0055] The communication processing unit 14 determines that the connection has been rejected when the reject button on the inquiry screen is indicated by an operation signal input from the input device 32, or when the allow button is not indicated even after a predetermined waiting time has elapsed since the display of the inquiry screen. The communication processing unit 14 then sends a rejection response to the remote device that sent the connection request. At this point, the local device fails to establish a connection with the remote device.

[0056] The output processing unit 16 may identify the remote device and display an operation screen on the display 24 to instruct the start of communication. At this time, the communication processing unit 14 may receive an operation signal from the input device 32 indicating a connection request to the remote device. When the communication processing unit 14 receives an operation signal indicating a connection request from the input device 32, it sends the connection request to the remote device indicated by the operation signal. When the communication processing unit 14 receives an acknowledgment from the remote device in response to the connection request, it determines that the connection with the remote device is permitted. At this time, a connection is established between the local device and the remote device. When the communication processing unit 14 receives a rejection response from the remote device in response to the connection request, or when it does not receive an acknowledgment even after a predetermined waiting time has elapsed since sending the connection request, it determines that the connection with the remote device is not permitted. At this time, the local device does not establish a connection with the remote device.

[0057] The communication processing unit 14 sends and receives communication data with the remote device once a connection with the remote device has been established. The communication processing unit 14 transmits the multiplexed data input from the video processing unit 12 to the receiving device. The communication processing unit 14 may include either or both of the following in the multiplexed data: audio data input from the audio system 26, or application screen data showing a screen obtained by executing another application program (sometimes referred to as an "app" in this application) (sometimes referred to as an "app screen" in this application), and transmit this to the receiving device. The app may provide functions such as a chat function or a document display function.

[0058] When the communication processing unit 14 receives multiplexed data from the other device, it outputs the received multiplexed data to the synthesis unit 126. If the received multiplexed data contains multiplexed audio data, the communication processing unit 14 separates the audio data from the multiplexed data and outputs the separated audio data to the audio system 26 via the output processing unit 16. If the received multiplexed data contains multiplexed application screen data, the communication processing unit 14 separates the application screen data from the multiplexed data and outputs the separated application screen data to the output processing unit 16.

[0059] The output processing unit 16 performs processing to output various information related to communication with the other device. The output processing unit 16 configures a display screen that includes the reconstructed image shown in the video data input from the synthesis unit 126. The output processing unit 16 outputs display data showing the configured display screen to the display 24. The display 24 displays the display screen including the reconstructed image based on the display data.

[0060] When application screen data is input from the communication processing unit 14, the output processing unit 16 may configure a display screen that includes the application screen shown in the application screen data. When audio data is input from the communication processing unit 14, the output processing unit 16 may output the input audio data to the audio system 26 and have the audio played from the speaker. The output processing unit 16 may output display data showing various setting screens to the display 24.

[0061] Next, the video communication processing according to this embodiment will be described. Figure 4 is a flowchart illustrating the video communication processing according to this embodiment. The video communication processing illustrated in Figure 4 shows a series of processes from capturing the original video in terminal device 1-1 to displaying the restored video in terminal device 1-2. Furthermore, the following description mainly assumes that there are two stages in the elemental video. The stage with the lowest importance may be called the "lowest stage," the subject or its image related to the lowest stage may be called the "non-essential part," and the video showing the non-essential part may be called the "non-essential part video." Also, a stage that is more important than the lowest stage may be called the "high stage," the subject or its image related to the high stage may be called the "essential part," and the video showing the essential part may be called the "essential part video." The number of key image segments detected from the original video is not necessarily limited to one; it may be two or more, or it may be zero.

[0062] Terminal device 1-1 executes the processes in steps S102 to S108. (Step S102) Camera 25 captures images. (Step S104) The analysis unit 122 uses the video captured by the camera 25 as the source video and uses a predetermined model to detect elemental video from the source video that contains a specific type of subject as its main component. The analysis unit 122 also detects feature points from the source video that indicate predetermined morphological characteristics of the subject. (Step S106) The analysis unit 122 uses a model separate from the detection of essential parts to detect elemental images representing non-essential parts from the original image. (Step S108) The analysis unit 122 increases the information density of the essential part video showing essential parts compared to the non-essential part video showing non-essential parts, and the communication processing unit 14 transmits the multiplexed data, which is multiplexed with the encoded data of the essential part video and the non-essential part video and feature information showing feature points, to the terminal device 1-2 using the network NW.

[0063] Terminal device 1-2 executes the processes in steps S110 to S116. (Step S110) The communication processing unit 14 receives multiplexed data from the terminal device 1-1. The synthesis unit 126 separates the encoded data of the essential part of the video and the non-essential part of the video, as well as feature information indicating feature points, from the multiplexed data. (Step S112) The synthesis unit 126 adjusts the information density of the non-essential video so that it is equal to the information density of the essential video. (Step S114) The synthesis unit 126 reconstructs the restored image by superimposing the essential image so that the positions of the feature points in the essential image are positioned at the positions of the feature points indicated in the feature information in the adjusted non-essential image. (Step S116) The output processing unit 16 displays the reconstructed restored video on the display 24. After that, the process shown in Figure 4 is terminated.

[0064] In addition, in communication system S1, video may also be transmitted from terminal device 1-2 to terminal device 1-1. In that case, terminal device 1-2 further performs the processing in steps S102 to S108, and terminal device 1-1 performs the processing in steps S110 to S116.

[0065] Next, we will show an example of video or image processing. Figure 5 illustrates a still image that makes up the original video. The still image illustrated in Figure 5 is captured by the camera 25 of terminal device 1-1 during communication with terminal device 1-2. This still image shows the front view of the user's head and chest at a certain time. Figure 6 illustrates still images that constitute the essential parts of the image. The example still images show the left eye, right eye, and lips of the user's face, respectively, as essential parts detected by the analysis unit 122 from the still image in Figure 5. The left eye, right eye, and lips are organs whose shape and position change significantly in response to the user's emotions. These organs tend to attract attention from the other party during communication. In the essential parts image, unlike the non-essential parts image, the reduction in information density is prevented or mitigated, so communication is not impaired.

[0066] Figure 7 shows an example of feature information detection from the original image. The feature information shown includes feature points and edges detected from the still image exemplified in Figure 5. Feature points, such as the outer corners of both eyes, the corners of the mouth, and the cleft of the mouth, are indicated by "x" marks. The outer edge of the bottom of the face, including the chin, is shown as a curve. Figure 8 illustrates still images that constitute non-essential video. The exemplified still images are detected from the original video by the analysis unit 122 and have a lower resolution than the essential video. The exemplified still images mainly represent the background, which consists of windows. Since the background is not of interest to the user of the other device, a lower information density is acceptable in the non-essential video showing the background compared to the essential video.

[0067] Figure 9 illustrates still images that make up the reconstructed image. The reconstructed image shown is obtained by superimposing the essential image onto the non-essential image in the synthesis unit 126. The essential image is superimposed onto the non-essential image such that the feature points of the essential image are positioned at the locations indicated by the feature information.

[0068] Next, a specific example of combining essential and non-essential video will be explained. Figure 10 is an explanatory diagram illustrating the combination of essential and non-essential video. However, the example shows the case where the frame rate of the essential video is transmitted at three times the frame rate of the non-essential video. In the combining unit 126, still images for each frame constituting the non-essential video are interpolated so that the frame rate becomes three times. When interpolating still images, the combining unit 126 may repeat each still image constituting the non-essential video three times with the same frame period as the element video, or it may generate it by morphing using the still images for each frame. Figures 10(a) to (d) show the non-essential video and essential video for each frame from time t=t0 to t0+3ΔT (=t0+δT; ΔT and δT represent the frame intervals of the essential video and non-essential video, respectively) as thin line drawings and thick line drawings, respectively. The edges that form the outline of the essential video are used as morphological features. The still image constituting the main part of the image is superimposed on the still image constituting the non-elemental part of the image, where its outline is positioned as indicated by the feature information. Figures 10(a) to (d) show that the subject appearing in the main part of the image moves sequentially to the right from time t=t0 to t0+3ΔT. Therefore, even if the frame rates of the main part of the image and the non-main part of the image are different, the movement of the subject shown in the main part of the image is smoothly reproduced on the non-main part of the image, which has an adjusted frame rate.

[0069] Furthermore, the essential parts uniformly determined based on the type of subject do not necessarily contribute to information transmission between users. Therefore, the analysis unit 122 may determine subjects whose movement amount between frames exceeds a predetermined reference value as non-essential parts. This is because, for subjects moving at high speed, a decrease in information density is less likely to be perceived as a decrease in image quality, and they do not contribute to the transmission of precise information. In addition, the analysis unit 122 may exclude subjects whose size falls outside a predetermined range from essential parts, and determine subjects whose size falls within the predetermined range as candidates for essential parts.

[0070] Next, an example of the essential part determination process will be explained using Figure 11. Figure 11 is a flowchart illustrating the essential part determination process according to this embodiment. (Step S202) The analysis unit 122 identifies the subject appearing in the original video acquired from the camera 25. If the identified subject is of a predetermined type (Step S202 YES), the analysis unit 122 proceeds to the process in Step S204. If the identified subject is of a different type than the predetermined type (Step S202 NO), the display area of ​​that subject is determined to be non-essential, and the process in Figure 11 is terminated.

[0071] (Step S204) The analysis unit 122 determines that the display area of ​​a predetermined type of subject is non-essential if the amount of movement from the subject appearing in the immediately preceding frame is greater than or equal to a predetermined reference amount (Step S204 YES), and terminates the process shown in Figure 11. If the amount of movement is less than the predetermined reference amount (Step S204 NO), the process proceeds to Step S206.

[0072] (Step S206) The analysis unit 122 determines whether the size of a predetermined type of subject is within a predetermined size range (for example, 1 / 32 to 1 / 64 of one side of the display screen). If the size is within the predetermined range (Step S206 YES), the analysis unit 122 determines that the image of the subject is the essential part. After that, the process shown in Figure 11 is terminated. If the size is smaller than or larger than the predetermined range (Step S206 NO), the analysis unit 122 determines that the display area of ​​the subject is not essential and terminates the process shown in Figure 11.

[0073] As explained at the beginning, transmitting high-resolution video directly requires a large transmission capacity (i.e., bandwidth). When transmitting 4K video (3840 pixels x 2160 pixels) at 30 frames per second based on the method specified in ITU-T H.264, the required communication capacity is 16 Mbps. However, in this embodiment, the 4K video is divided into essential and non-essential parts as the original video, and the information density of the essential parts is not reduced from the original video, while the information density of the non-essential parts is reduced from the original video. This reduces the overall communication capacity of the video without drastically degrading subjective quality. For example, when the number of pixels in the essential parts, which constitute a portion of the area of ​​one frame, is set to 540 pixels x 960 pixels at 30 frames per second, and the number of pixels in the non-essential parts, set to 240 pixels x 135 pixels at 10 frames per second, and the required transmission capacity is 1 Mbps for the essential parts and 0.2 Mbps for the non-essential parts. Overall, the total speed is 1.2 Mbps, which is less than 1 / 12 the transmission capacity required to transmit 4K video directly.

[0074] The above explanation primarily uses the example of a communication system S1 with two connections and involving the transmission of video data between two terminal devices 1, but it is not limited to this. The communication system S1 may have three or more connections and be applied to communication between three or more terminal devices 1. In that case, at least one of the three or more terminal devices 1 is designated as the first device, and the other multiple terminal devices are designated as the second devices. The first device transmits common multiplexed data based on the video acquired by its own device to each of the second devices. The first device receives multiplexed data based on the video acquired by each individual device from each of the second devices. The first device is capable of synthesizing the reconstructed video based on the respective multiplexed data. The first device may present all of the synthesized reconstructed video in parallel, or it may present only a part of the reconstructed video. The communication processing unit 14 can identify three or more terminal devices 1 participating in a single communication using the above method.

[0075] The analysis unit 122 may reduce the amount of information in the multiplexed data per transmitting device as the number of connections participating in a single communication increases. Here, the analysis unit 122 may reduce the information density of non-essential video or reduce the information density of essential video as the number of connections increases. Multiple stages of candidate subject types for essential parts are set in advance for the analysis unit 122, and the range of subjects that can be candidates for essential parts may be narrowed as the number of connections increases. When the number of connections exceeds a predetermined upper limit for the number of connections, the analysis unit 122 may not perform essential part analysis, but instead convert the entire original video into non-essential video and stop transmitting the essential video.

[0076] The communication processing unit 14 may determine whether or not speech is being uttered using known speech processing techniques based on the audio data input to it. The communication processing unit 14 may transmit multiplexed data based on the video acquired during the speech period in which speech is determined to be occurring, and may stop transmitting multiplexed data based on the video acquired during the non-speech period in which speech is not determined to be occurring. In addition, the communication processing unit 14 may treat the entire original video as non-essential video and stop transmitting the essential video during the speech period.

[0077] As described above, the communication system S1 according to this embodiment comprises at least a first device (e.g., terminal device 1-1) and a second device (e.g., terminal device 1-2). The first device comprises a first video processing unit (e.g., analysis unit 122 of video processing unit 12) that divides the original video into multiple stages of elemental video of different importance based on the appearance of the subject, and adjusts the information density of the elemental video so that the higher the importance of the stage, the higher the relative information density of the elemental video; and a first communication processing unit (e.g., communication processing unit 14) that transmits the multiple stages of elemental video to the second device. The second device comprises a second communication processing unit (e.g., communication processing unit 14) that receives the multiple stages of elemental video from the first device; and a second video processing unit (e.g., synthesis unit 126 of video processing unit 12) that synthesizes the multiple stages of elemental video to reconstruct a restored video and outputs the restored video. Certain types of subjects may include the upper body of a person (e.g., head, hands, arms, etc.) or a display medium that displays content (e.g., touch panel, bulletin board, etc.). In this configuration, elemental video footage with reduced information density according to importance is transmitted from the first device to the second device, and a reconstructed video is output by combining elemental video footage in multiple stages. While reducing the overall communication capacity required for the communication system S1, it is possible to ensure quality by relatively increasing the information density for subjects important for communication, thereby maintaining smooth communication between the first and second devices through video communication.

[0078] The first video processing unit may detect feature information indicating predetermined morphological features (e.g., feature points, edges) and positions of the subject from the original video, and the second video processing unit may reconstruct the restored video by superimposing multiple stages of elemental video so that the morphology is positioned at the locations indicated by the feature information. In this configuration, the position where elemental images, which constitute part of the video to be transmitted, should be superimposed is identified using morphological features indicated by feature information as a clue. Furthermore, processing efficiency can be improved by omitting the process of analyzing feature information in the second video processing unit.

[0079] The second video processing unit may convert the multi-stage elemental video so that its information density is equal to the highest information density among the multi-stage elemental video, and then reconstruct the restored video by superimposing the converted multi-stage elemental video. In this configuration, the information density is standardized among the multiple stages of elemental video so that it is equal to the highest information density among the multiple stages of elemental video. Therefore, the processing load related to the superposition of multiple stages of elemental video can be reduced, and the quality of the restored video can be ensured.

[0080] The second video processing unit may reconstruct the restored image by superimposing multiple stages of elemental images, prioritizing those with higher importance. With this configuration, even if elemental images overlap across multiple stages, the elemental images from the stages with higher importance will be reflected in the restored image. Furthermore, it reduces the quality degradation caused by information loss due to the overlapping of elemental images.

[0081] The first video processing unit detects the amount of movement of a subject in a specific manner from the original video. If the amount of movement is outside a predetermined range, it determines that the element video representing the subject is an element video of the lowest level, which is the least important level. If the amount of movement is within a predetermined range, it may determine that the element video representing the subject is a candidate for an element video of a higher level, which is more important than the lowest level element video. In this configuration, among subjects of a specific type, images of subjects whose movement is within a predetermined range become candidates for high-level elemental images, while images of subjects whose movement exceeds the predetermined range become low-level elemental images. By setting the information density of images of subjects whose movement exceeds the predetermined range and do not contribute to communication to the lowest level, transmission capacity can be reduced without impairing communication.

[0082] The first video processing unit detects the size of an image of a subject of a specific nature from the original video. If the size is outside a predetermined range, it determines that the element video representing the subject is a basic video of the lowest level, which is the least important level. If the size is within the predetermined range, it determines that the element video representing the subject is a candidate for a high-level element video, which is more important than the lowest-level element video. In this configuration, images of subjects within a specific size range become candidates for high-level elemental images, while images of subjects exceeding a predetermined size range become low-level elemental images. By setting the information density of images of subjects exceeding a predetermined size range that do not contribute to communication to the lowest level, transmission capacity can be reduced without impairing communication.

[0083] Although embodiments of the present invention have been described in detail above with reference to the drawings, the specific configurations are not limited to the embodiments described above, and include designs and the like that do not depart from the spirit of this invention. The configurations described in the embodiments described above can be combined in any way. [Explanation of Symbols]

[0084] S1...Communication system, 1(1-1, 1-2)...Terminal device, 10...Host system, 12...Video processing unit, 14...Communication processing unit, 16...Output processing unit, 22...ROM, 23...Auxiliary storage device, 24...Display, 25...Camera, 26...Audio system, 27...Communication module, 28...Input / Output I / F, 31...EC, 32...Input device, 33...Power supply circuit, 36...Power switch, 122...Analysis unit, 124...Multiplexing unit, 126...Synthesis unit, NW...Network

Claims

1. A communication system comprising at least a first device and a second device, The first apparatus is The original video is divided into multiple levels of elemental video with varying degrees of importance based on the subject's characteristics. A first video processing unit adjusts the information density of elemental video so that it becomes relatively higher for stages of higher importance, The system includes a first communication processing unit that transmits the multiple-stage elemental images to the second device, The second device is A second communication processing unit that receives the multiple-stage elemental images from the first device, The system includes a second video processing unit that synthesizes multiple stages of the aforementioned elemental video to reconstruct a restored video. Communication system.

2. The first video processing unit is: From the original video footage, characteristic information indicating the morphological features and position of the subject is detected. The second video processing unit is: The reconstructed image is reconstructed by superimposing multiple stages of the elemental images so that the morphological features are positioned at the locations indicated in the feature information. The communication system according to claim 1.

3. The second video processing unit converts the multiple-stage elemental video so that its information density is equal to the highest information density among the multiple-stage elemental video, and then superimposes the converted multiple-stage elemental video to reconstruct the restored video. The communication system according to claim 2.

4. The second video processing unit reconstructs the restored video by superimposing multiple stages of elemental video so that elements with higher importance are given priority. The communication system according to claim 3.

5. The first video processing unit is: The amount of movement of a subject in a specific manner is detected from the original video footage. If the amount of movement is outside the predetermined range of movement, the element image showing the subject is determined to be the lowest-level element image, which is the least important level. If the amount of movement is within the range of the amount of movement, the element image showing the subject is determined to be a candidate for a higher-level element image, which is of higher importance than the lowest-level element image. The communication system according to claim 3.

6. The first video processing unit is: The size of the image of a subject in a specific manner is detected from the original video footage. If the aforementioned size falls outside the predetermined size range, the element image showing the subject is determined to be the lowest-level basic image, which is the least important level. If the size is within the range of the size, the element image showing the subject is determined to be a candidate for a higher-level element image, which is of higher importance than the lowest-level element image. The communication system according to claim 3.

7. The subject in the aforementioned specific form is the upper body of a person, or a display medium that displays content. The communication system according to claim 5 or claim 6.

8. The original video is divided into a first stage of multiple elemental video segments of varying importance based on the subject's characteristics. A video processing unit adjusts the information density of elemental video so that it increases with increasing importance, The system includes a communication processing unit that transmits the first multi-stage elemental video to another device, The aforementioned communication processing unit, Receiving at least two sets of elemental images from the aforementioned other device, The aforementioned video processing unit, The second multi-stage elemental video is synthesized to reconstruct the restored video. Communication device.

9. In a communication system comprising at least a first device and a second device, The first apparatus is The original video is divided into multiple levels of elemental video with varying degrees of importance based on the subject's characteristics, The steps include adjusting the information density of the elemental video so that it becomes relatively higher for stages of higher importance, and The steps of transmitting the aforementioned multi-stage elemental images to the second device are performed, The second device is The steps include receiving the multiple-stage elemental images from the first device, The process involves performing the steps of: synthesizing multiple stages of the aforementioned elemental images to reconstruct a restored image; Communication method.