Remote virtual interaction device and system based on multiple capturers
By directly transmitting image and sound information through a multi-capture device, the problem of insufficient realism in existing remote interaction technologies is solved, achieving a low-cost and efficient remote interaction effect.
Patent Information
- Application Number
- CN202420060027.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2034-01-09
AI Technical Summary
Existing remote interaction technologies lack realism, have high equipment costs and require large amounts of computation, making it difficult to meet the needs of instant communication.
The device employs a multi-capture device, including an image capturer and an audio capturer, to transmit image and audio information to the releaser via a network transmission device, directly presenting it to the interactive object and simulating face-to-face communication.
It enables low-cost and efficient remote interaction, enhances realism, reduces equipment costs and computational complexity, and significantly improves the effectiveness of remote communication.
Smart Images

Figure CN223598202U_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of virtual interaction technology, in particular to a remote virtual interaction device and system based on multiple capturers. BACKGROUND
[0002] It is an important requirement for the modernization of human society to realize barrier-free remote interaction. The rapid development of modern information technology provides strong technical support for people's remote interaction. With the help of computer network communication technology and various client programs, people can communicate remotely through short messages, telephone calls, WeChat, and other ways, such as text chatting, voice calls, video calls, and other ways. Among these interaction methods, text chatting appeared earlier. Although it can achieve information transmission, it lacks auditory experience and is gradually replaced by voice calls in some situations. Voice calls can meet the need for hearing, but lack visual perception. Therefore, video calls, which are widely welcomed, combine visual and auditory effects and can fully display the current situation of both parties in communication. Despite this, people's pursuit does not stop there. Remote video calls still lack enough authenticity compared to face-to-face communication, and the ideal of barrier-free remote interaction needs to be further improved.
[0003] In order to make remote interaction also achieve the effect of face-to-face communication, people have made continuous efforts. One feasible way is to use sensors to scan parts such as the human eye or face, and then load the scanned information into a pre-established three-dimensional character model, so as to realize real-time three-dimensional display of both parties in communication. Specifically, first, a local computer technology is used to scan a person to be remotely interacted with to establish a three-dimensional model corresponding to the person; then both parties in communication wear eye devices based on virtual reality technology, use multiple sensors integrated on the devices to scan eye movement and facial changes, capture signal data such as eye rotation, facial changes, and body movements in real time, and load these signal data into the pre-established three-dimensional character model to drive the three-dimensional character model in the virtual scene to keep synchronization with the person in reality, so that the receiving party sees the actions and expressions of the virtual character, as if the real person's performance is reproduced, achieving the effect of augmented reality communication.
[0004] However, as can be seen from the foregoing description, this way virtualizes the person and establishes a three-dimensional character model, which takes tens of minutes at the shortest and several hours at the longest, making it difficult to meet the needs of real-time communication in reality. Moreover, in order to synchronize the actions and expressions of the real person, a large number of sensors must be installed to collect data at points such as facial parts, eye rotation, and body movements, and after collection, data synthesis calculation must be performed to effectively drive the three-dimensional character model, which inevitably leads to high equipment cost, large amount of calculation, and long time consumption, and cannot meet the goal of barrier-free communication and interaction of people. CONTENT OF THE UTILITY MODEL
[0005] The embodiment of the present application provides a multi-captor-based remote virtual interaction device and system, which is used for enhancing the reality of remote interaction, reducing calculation, saving cost and improving remote communication interaction effect.
[0006] In one aspect, the multi-captor-based remote virtual interaction device provided by the embodiment of the present application comprises:
[0007] The first type of captor comprises at least one group of image captors, each group of image captors comprising a first image captor and a second image captor, and is used for capturing image information of a predetermined position of an interaction object on the side where the first type of captor is located, wherein the lateral distance between the first image captor and the second image captor is a first predetermined distance, and the first predetermined distance is within the eye distance range of the interaction object on the opposite side of the side where the first type of captor is located.
[0008] The second type of captor comprises at least one sound captor, and is used for capturing sound information at the position of an interaction object on the side where the second type of captor is located.
[0009] The first type of release comprises at least one group of image releases, each group of image releases comprising a first image release and a second image release, and is used for presenting the image information captured by the first image captor and the second image captor on the opposite side of the side where the first type of release is located, respectively, in the left and right eye view ranges of an interaction object on the side where the first type of release is located.
[0010] The second type of release comprises at least one sound release, and is used for releasing the sound information captured by the second type of captor on the opposite side of the side where the second type of release is located to an interaction object on the side where the second type of release is located.
[0011] Preferably, the first image captor and the second image captor are arranged within a predetermined pitch angle range of a horizontal view plane in front of the interaction object on the side where the first type of captor is located.
[0012] Preferably, the vertical distance between the first image captor and the second image captor relative to the interaction object on the side where the first type of captor is located is a second predetermined distance, and the second predetermined distance is within the social distance range of the real interaction of the interaction object.
[0013] Preferably, the sound release is arranged at a third predetermined distance below the middle part of the lateral distance between the first image captor and the second image captor, and the third predetermined distance is within the distance range between the eyes and the mouth of the interaction object on the opposite side of the side where the first type of captor is located.
[0014] In another aspect, the multi-captor-based remote virtual interaction device provided by the embodiment of the present application comprises:
[0015] The first type of capture device includes at least one set of image capture devices, each set of image capture devices including a first image capture device and a second image capture device, configured to capture image information including a predetermined position of an interactive object on a side where the first type of capture device is located, so as to be transmitted to a corresponding image release device in a first type of release device on a side opposite to the side where the first type of capture device is located via a network transmission device, a lateral distance between the first image capture device and the second image capture device being a first predetermined distance, the first predetermined distance being within an eye distance range of the interactive object on the side where the first type of capture device is located;
[0016] The second type of capture device includes at least one sound capture device, configured to capture sound information at a position of an interactive object on a side where the second type of capture device is located, so as to be transmitted to a corresponding sound release device in a second type of release device on a side opposite to the side where the second type of capture device is located via a network transmission device.
[0017] In another aspect, the embodiment of the present application provides a remote virtual interaction device based on a plurality of capture devices, which includes:
[0018] The first type of release device includes at least one set of image release devices, each set of image release devices including a first image release device and a second image release device, configured to present image information of an interactive object captured by the first image capture device and the second image capture device respectively to the interactive object on a side where the first type of release device is located within a left eye and right eye visual range of the interactive object on the side where the first type of release device is located;
[0019] The second type of release device includes at least one sound release device, configured to release sound information captured by the second type of capture device to an interactive object on a side where the second type of release device is located.
[0020] In another aspect, the embodiment of the present application provides a remote virtual interaction system based on a plurality of capture devices, which includes:
[0021] The first type of capture device on a remote interaction side includes at least one set of image capture devices, each set of image capture devices including a first image capture device and a second image capture device, configured to capture image information including a predetermined position of an interactive object on a side where the first type of capture device is located, a lateral distance between the first image capture device and the second image capture device being a first predetermined distance, the first predetermined distance being within an eye distance range of the interactive object on the side where the first type of capture device is located;
[0022] The second type of capture device on the remote interaction side includes at least one sound capture device, configured to capture sound information at a position of an interactive object on a side where the second type of capture device is located;
[0023] The first type of release device located on the other side of the remote interaction includes at least one set of image release devices, each set of image release devices including a first image release device and a second image release device, for respectively presenting image information captured by the first image capture device and the second image capture device located on the opposite side of the first type of release device to the left and right eye visual ranges of the interactive object located on the side of the first type of release device;
[0024] The second type of release device located on the other side of the remote interaction includes at least one sound release device, for releasing sound information captured by the second type of capture device located on the opposite side of the second type of release device to the interactive object located on the side of the second type of release device;
[0025] The network transmission device is configured to transmit image information captured by the first type of capture device at a predetermined position of the interactive object and sound information captured by the second type of capture device at a position of the interactive object to the first type of release device and the second type of release device.
[0026] In another aspect, the embodiment of the present application provides a remote virtual interaction system based on multiple capture devices, which includes:
[0027] The first type of capture device includes at least one set of image capture devices, each set of image capture devices including a first image capture device and a second image capture device, for capturing image information including a predetermined position of an interactive object located on the side of the first type of capture device, a lateral distance between the first image capture device and the second image capture device being a first predetermined distance, and the first predetermined distance being within an eye distance range of the interactive object located on the side of the first type of capture device;
[0028] The second type of capture device includes at least one sound capture device, for capturing sound information including a position of an interactive object located on the side of the second type of capture device;
[0029] The first type of release device includes at least one set of image release devices, each set of image release devices including a first image release device and a second image release device, for respectively presenting image information captured by the first image capture device and the second image capture device located on the opposite side of the first type of release device to the left and right eye visual ranges of the interactive object located on the side of the first type of release device;
[0030] The second type of release device includes at least one sound release device, for releasing sound information captured by the second type of capture device located on the opposite side of the second type of release device to the interactive object located on the side of the second type of release device;
[0031] The network transmission device is configured to transmit image information captured by the first type of capture device including a predetermined position of an interactive object located on the side of the first type of capture device and sound information captured by the second type of capture device to the first type of release device and the second type of release device located on the opposite side of the first type of capture device.
[0032] Preferably, the first image capture, the second image capture is a camera, and / or, the sound capture in the second type of capture is a microphone, the sound release in the second type of release is a loudspeaker, and / or, the image release in the first type of release is a three-dimensional display device.
[0033] Preferably, the three-dimensional display device is AR / VR glasses, and / or naked-eye 3D display screen, and / or raster 3D perspective glasses.
[0034] Compared with the prior art, the embodiment of the present application no longer needs to model the three-dimensional character model in advance, and no longer needs a large number of cameras to capture information such as eyeball rotation, facial expression, and body change of the interactive object, and no longer needs complex comprehensive data calculation to drive the three-dimensional character model. This low-cost, high-efficiency, and more realistic remote interaction mode eliminates or weakens the obstacles of remote interaction, and significantly improves the effect of remote interaction. BRIEF DESCRIPTION OF DRAWINGS
[0035] The accompanying drawings, which are included to provide a further understanding of the present application, form a part of the present application and illustrate the illustrative embodiments of the present application and together with the description serve to explain the present application. In the drawings:
[0036] Figure 1a A scene example diagram of the embodiment of the present application;
[0037] Figure 1b A schematic diagram of the framework structure of the embodiment of the present application;
[0038] Figure 2 A schematic diagram of the lateral spacing range between the image captures of the embodiment of the present application;
[0039] Figure 3 A schematic diagram of the pitch relationship between the image captures and the interactive object of the embodiment of the present application;
[0040] Figure 4 A schematic diagram of the vertical spacing range between the image captures and the interactive object of the embodiment of the present application;
[0041] Figure 5 A schematic diagram of the relationship between the image captures and the sound release (sound capture) of the embodiment of the present application;
[0042] Figure 6a A scene diagram of the embodiment of the present application;
[0043] Figure 6b The embodiment of the present application Figure 6a The effect diagram under the scene shown. DETAILED DESCRIPTION
[0044] Before various embodiments of the present application are fully described, some basic background and basic terminology concepts are briefly introduced for ease of understanding. With the development of information computer technology, people have been able to easily implement remote communication to solve the needs of two or more people to interact at different locations. However, as mentioned in the background section, the current remote interaction mode, whether it is text chat, voice call, or video call, has the problem of "low sense of reality", and even if it can solve the problem of "reality" to some extent, it needs to spend a high cost, which is not conducive to the popularization of information interaction technology. The so-called "sense of reality" here refers to the use of real interaction effect to measure remote interaction effect, that is, making the interaction between two parties thousands of miles apart feel like face-to-face communication through technological innovation. The closer to the real interaction effect, the lower the cost, the more it is the direction of people's technological efforts. The reason why the remote communication effect in the prior art is lower than the real interaction effect is that when face-to-face interaction, the interactive object seen by the human eye is three-dimensional and all-around, while in remote communication, the remote end interactive personnel sees the local end interactive personnel, which is always flat. Even if a three-dimensional image is presented on the remote end interactive device, it is a result of computer software simulation, not a "three-dimensional image" actually seen by the eyes of the remote end interactive personnel, and thus lacks reality. Based on this background, it is still necessary to better utilize the visual characteristics of human beings on the basis of existing technology to improve or enhance the reality of interaction, overcome or weaken the communication defects brought by non-face-to-face interaction.
[0045] In the introduction of the above background, the term "interaction" is mentioned. First, the concept of "interaction" at least covers two interaction subjects. In the embodiments of the present application, they are called "interaction objects". In order to distinguish different interaction subjects, the whole interaction chain can be viewed from the perspective of one of the "interaction objects", that is, calling itself a local side interaction object and calling the other party a remote side interaction object. As can be seen, this is a mutual call. From the perspective of one of the interaction objects, it is the local side and the other party is the remote side. However, from the perspective of the other interaction object, the object is the remote side and itself is the local side. Of course, in addition to the subjective call from the perspective of the self, the interaction objects in different positions can also be objectively distinguished by taking the device as a reference. For example, a first type of capturer is arranged on the local side and a first type of releaser is arranged on the opposite side. Then, the interaction object on the local side can be called the interaction object on the first type of capturer side and the interaction object on the remote side can be called the object on the first type of releaser side. This distinction of the interaction object from the perspective of the self or the objective reference is particularly important in a symmetrical system (for example, the devices required on both ends of the interaction are the same). Second, in the embodiments of the present application, attention should also be paid to the difference between "interaction" and "face-to-face interaction". In the process of remote interaction, in addition to the interaction subjects, there can also be an interaction medium. The interaction medium can be a network device system, which can transmit the situation of the local side "interaction object" to the "interaction object" on the remote side and vice versa. This interaction by means of network technology devices is remote interaction and is also virtual interaction. Therefore, in some cases in the embodiments of the present application, the terms "interaction", "remote interaction" and "remote virtual interaction" all express the meaning of "interaction". Third, in the concept of "interaction", in terms of its essence, it also embodies an "initiation-response" mechanism, that is, there is an initiator of an event and there is also a responder to the event. In the embodiments of the present application, the interaction can be one-way or two-way. In terms of one-way, for example, the event initiator sends an information to the event acceptor, the event acceptor receives the information and gives a confirmation, which can be considered as an "interaction". In terms of two-way, after the event acceptor responds to the event initiated by the event initiator, the event acceptor can also initiate a new event as an initiator, and the previous initiator becomes a responder to the newly initiated event in this session. This "initiation-response" mechanism is the essential meaning of the interaction. Thus, it can be distinguished from a simple "monitoring" behavior. For example, a monitoring camera also remotely monitors the picture in a predetermined range through a network transmission device and stores the picture in the form of a continuous data stream. However, the subject that receives the picture does not "interfere", "influence" or "participate" in the formation of the monitoring picture.Of course, these two are possible to transform, especially from simple monitoring into monitoring the interaction between the two ends of the device, such as, in the acceptance of the picture this paragraph, can be monitored in the picture of the person or animal shouting, that is, the voice is presented to the monitored range.
[0046] The above description of "interaction" is to fundamentally understand the concept of the present application. The present application will also involve the capture device, the release device or the like. Obviously, this naming method is from the functional point of view, the image capture device is used to capture the video picture in its recording range, and the release device is used to restore the image or sound transmitted from the remote end to the local end. Any device or device that realizes this function can be used to achieve the purpose of the present application.
[0047] In the process of solving the problems existing in the prior art, the embodiment of the present application proposes the following technical solutions. The following will be described in detail in combination with the introduction of the figures and the foregoing related terms. Referring to FIG. 1, wherein: Figure 1a is a scene example diagram of the embodiment of the present application, Figure 1b is a framework structure diagram of the embodiment of the present application, and the exemplary device composition and layout of the remote virtual interaction device are embodied in the scene example diagram and the framework structure diagram.
[0048] The device includes two types of capture devices: one type is a capture device (U11, U12) for capturing at least image information, and the other type is a capture device (U2) for capturing at least sound information. It is worth noting that the classification here is not a "complete classification", but only a need for convenient calling. Similarly, the "class" here is not only the two classes of image and sound. In a class, other information such as smell, light, geographical location, etc. can be captured in addition to image information and sound information. Therefore, the "multiple" in the invention name in the embodiment of the present application can be understood in two levels. One is in the "class" level, and the capture device can include two or more types of capture devices. The other is in the level below the "class", and each type of capture device can include two or more image capture devices, one or more sound capture devices, or capture devices that capture other information. Regardless of which level, the device of the embodiment of the present application is a "multiple" capture device, and is especially suitable for "multiple-to-multiple" remote interaction scenarios.
[0049] The first type of capture device includes at least one set of image capture devices, each set of image capture devices including at least two image capture devices. In accordance with the foregoing description, the at least one set of "image" capture devices indicates that, in addition to capturing images, other information can also be captured as needed in the first type of capture device, and appropriate capture devices or equipment can be provided for these information. As shown in FIG. 1, when two image capture devices are included, they are referred to as a first image capture device U11 and a second image capture device U12 for ease of distinction. Both of the image capture devices can capture image information including a predetermined position of an interactive object on the side of the capture device, such as the eyes, face, upper body, etc. of the interactive object, and of course, can further include image information of the surrounding environment including the interactive object. The specific range of capture depends on the range of the side interactive object that the opposite side interactive object wishes or is allowed to see, and the layout of the first image capture device and the second image capture device is determined according to the "wished or allowed range". The layout includes at least the position arrangement between the first image capture device and the second image capture device, and the position arrangement between the first image capture device and the second image capture device and the interactive object. In order to more clearly show the effect of "face-to-face" interaction in reality, since the capture device is mainly used to capture the face of the interactive object, and more importantly, to capture specific parts (positions) such as the eyes of the interactive object, the lateral distance (the distance between the center points of the two image capture devices, since the lateral distance exists, the two image capture devices are generally not placed vertically overlapped) between the first image capture device and the second image capture device should be "linked" to the eye distance range of the interactive object. In general, see Figure 2 If the average value of the eye distance (the distance between the pupils of the eyes of a person) is τ (τ>0), the lateral distance range can be determined to be [τ-δ, τ+δ], where δ is a tolerance range under the factors of image definition, interactive reality, etc. Within this range, the image definition and / or interactive reality are acceptable. When the first image capture device and the second image capture device are constructed as an integral component / equipment (if not an integral component / equipment, the appropriate degree can be manually adjusted when the image capture device is placed), an adjustment knob can be provided between them, and the maximum range of the adjustment knob can be determined as 2δ, so that the needs of different groups of people can be met in a larger range. The connection of the lateral distance between the two image capture devices and the eye distance range of the interactive object (the opposite side of the side where the first type of image capture device is located) is an important "trick" of the embodiments of the present application. This approach breaks the general sense that the image capture device is only used for recording images without considering the visual relationship with the interactive object. In fact, this approach is equivalent to extending the "eyes" of the interactive object on the opposite side, which is located far away on the opposite side, and suddenly appears in front of the local interactive object, achieving the effect of "thousand-mile eyes", and further enabling the local interactive object and the remote interactive object to communicate like "face-to-face", and maximizing the simulation of the real face-to-face interaction scene.
[0050] The second type of capture includes at least one sound capture U2, which is used to capture sound information at the location of the local interactive object. The sound information can be the sound emitted by the interactive object itself (such as speech, singing, humming, etc.), the sound emitted by the interactive object with the aid of other objects (such as percussion instrument sound, knocking on a table, etc.), or even the sound emitted by other objects at the location of the interactive object (such as a cell phone ring, a music player, a television, etc.), or other sounds that happen to enter the "interactive object" location. As can be seen, the "sound information" here is actually the sound at the local location, i.e., the sound associated with the interactive object. It is worth noting that at a certain moment, the sound information does not necessarily include the sound emitted by the interactive object itself. Here, "sound emitted by the interactive object" is considered from a long-term perspective: since the interaction is between two or more interactive objects at the local and remote locations, there should be sound emitted by the interactive object itself during the entire interaction, such as voice. The position of the sound capture U2 can be placed at any position in the environment of the local interactive object, as long as it is convenient for capturing sound. Typically, in order to facilitate sound collection, it can be placed at the chest or head of the interactive object, etc. In some embodiments, it is placed near the image capture, which is considered as how to more realistically simulate the "ear" of the remote interactive object, but it does not constitute a limitation. Figure 1a
[0051] In addition to the two types of captures, the device of the embodiments of the present application also provides two types of releases: one type is used to release at least the image captured by the first type of image capture, and the other type is used to release at least the sound captured by the second type of capture.
[0052] The "release" of the release is relative to the "capture" of the capture, i.e., the image released by the image release in the first type of release at the local location is relative to the image captured by the image capture in the first type of capture at the remote location, and the sound released by the sound release in the second type of release at the local location is relative to the sound captured by the sound capture in the second type of capture at the remote location. From the perspective of the remote location, there is also a corresponding relationship, i.e., the image released by the image release in the first type of release at the remote location is also relative to the image captured by the image capture in the first type of capture at the local location, and the sound released by the sound release in the second type of release at the remote location is also relative to the sound captured by the sound capture in the second type of capture at the local location.
[0053] More specifically, the first type of release can include at least one set of image release, and each set of image release can include a first image release U41 and a second image release U42, and the first image release can release the image captured by the first image capture on the remote side to one of the left and right eyes of the local side interactive object, and the second image release can release the image captured by the second image capture on the remote side to the other of the left and right eyes of the local side interactive object. Here, the release to the "eye" is actually the release into the line of sight range of the eye, so that the human eye can see the image. From the perspective of simulating a real scene, the image captured by the image capture located on the left side of the current interactive object (local side) should be released to the right eye of the remote side interactive object via the release on the remote side, and the image captured by the image capture located on the right side of the current interactive object should be released to the left eye of the remote side interactive object via the release on the remote side. Compared with the image release in the first type of release, the sound release in the second type of release is not so complex as the release of sound information, and it can directly release the sound information captured by the second type of capture (on the local side) via the second type of release (sound release U3).
[0054] After introducing the capture and release included in the device of the embodiments of the present application, it is necessary to further sort out the product forms that the embodiments of the present application can provide, which mainly include the following four cases (the first two cases are from the single side, and the last two cases are from the overall perspective):
[0055] The first case: single side symmetry. "Single side symmetry" means that the local side and the remote side are each a side, but the product forms are completely the same, such as Figure 1bThe local side or the remote side shown in the device components. For example, the local side includes a first image capture, a second image capture, a sound capture, and a first image release, a second image release, a sound release, and the remote side also includes a first image capture, a second image capture, a sound capture, and a first image release, a second image release, a sound release. The corresponding relationship here can be that the first image capture of the local side corresponds to the first image release of the remote side (of course, it can also correspond to the second image release, and the same below), the second image capture of the local side corresponds to the second image release of the remote side, the sound capture of the local side corresponds to the sound release of the remote side, and the first image release of the local side corresponds to the first image capture of the remote side (of course, it can also correspond to the second image capture, and the same below), the second image release of the local side corresponds to the second image capture of the remote side, and the sound release of the local side corresponds to the sound capture of the remote side. The real application scenario can be as follows: the local side and the remote side conduct a "one-to-one" meeting, and the person on one side (the local side or the remote side) needs to speak, and the person on the other side (the remote side or the local side) also needs to speak, so as to realize two-way interaction. When providing an actual product, a complete product can be provided for each single side, that is, the product of each side can be an integrated whole product containing part or all of the first image capture, the second image capture, the sound capture, and the first release and the second release, for example, the sound capture and the sound release are made into the same whole product. Of course, if each component is formed into an independent product, consumers can freely purchase and assemble or arrange the positions according to requirements.
[0056] The second case: single-sided asymmetric type. "Single-sided asymmetric" means that the local end and the remote end are each a side, but the product forms of the two single sides are not (completely) the same, such as Figure 1aThe local side and the remote side are shown. For example, the local side (remote side) only has the first image capture device, the second image capture device, and the sound capture device, the remote side (local side) only has the first image release device, the second image release device, and the sound release device, and the local side (remote side) does not have the first image release device, the second image release device, and the sound release device, and the remote side (local side) does not have the first image capture device, the second image capture device. Under this product form, the image information and the sound information are transmitted from the local side (remote side) to the remote side (local side), and the remote side (local side) responds to these information, but this response is not symmetrical. The real application scenario can be as follows: the local side and the remote side conduct a "one-to-many" conference, one person at the local side (remote side) speaks, and multiple people at the remote side (local side) pay attention to and listen to the speech of the person, thereby realizing two-way interaction. When providing an actual product, because it is asymmetrical, one side of the product provided is an integrated product containing the first image capture device, the second image capture device, and the sound capture device, and the other side of the product is an integrated product containing the first image release device, the second image release device, and the sound release device. Consumers can choose to purchase one or both of the two integrated products according to their own needs and main uses. Although the two products are asymmetrical, they should be matched when used, and the effective interaction of the embodiments of the present application is realized through the cooperation of the two products.
[0057] The third case is a bilateral (overall) symmetrical type. In this case, the product forms of the local side and the remote side are the same (as in the first case), and when viewed as a whole, the respective devices of the two sides together with the intermediate network device form a complete system symmetrical with respect to the network device. The network transmission device delivers the information about the interaction object captured by the first image capture device, the second image capture device, and the sound capture device of one side to the first image release device, the second image release device, and the sound release device of the opposite side, and delivers the information about the interaction object captured by the first image capture device, the second image capture device, and the sound capture device of the opposite side to the first image release device, the second image release device, and the sound release device of the same side, thereby realizing the interaction of the interaction object. This process is mutual, and the devices, transmission, and effects are symmetrical. When providing an actual product, a complete system can be prepared for consumers, and of course, the network device can also use the existing network architecture to connect the products of the local side and the remote side to realize interconnection, so that the local side, the network, and the remote side form a complete interactive system.
[0058] The fourth case is bilateral (overall) asymmetric. In this case, the product forms of the local side and the remote side are not (completely) the same (as in the second case), and when viewed as a whole, the respective devices of the two sides together with the intermediate network device form a complete system that is asymmetric with respect to the network device. The network transmission device mainly transmits the information about the interactive object captured by the first image capture device and the second image capture device and the sound capture device on one side to the first image release device, the second image release device and the sound release device on the opposite side, so as to realize the interaction between the interactive objects. This process is focused on one side and assisted by the other side, and is asymmetric in terms of devices, transmission and effects.
[0059] The above describes the structure (product form) of the embodiments of the present application in detail from various angles. It can be seen that the basic principle and superior technical effects of the present application are obtained. The present application not only focuses on the transmission and restoration of image, sound and other information of the two ends of remote interaction, but also focuses on how the devices, equipment or apparatuses that capture images and sounds simulate face-to-face interaction more realistically. Since the embodiments of the present application are provided with at least two image capture devices, and the lateral distance between the two capture devices is within the eye distance range of the interactive object on the opposite side, the images captured by the two image capture devices are respectively released to the left and right eyes of the interactive object on the opposite side by the image release device on the opposite side. At the same time, the sound information captured by the sound capture device is transmitted to the opposite side and released by the sound release device, which is equivalent to bringing the "eyes", "ears" and "mouth" of the interactive object on the opposite side close to the interactive object on the local side, thereby realistically simulating the scene of face-to-face communication, and realizing the "thousand-mile eyes", "wind-ear" and "big mouth" in technology.
[0060] Compared with the prior art introduced in the background art section, the embodiments of the present application no longer need to model the three-dimensional character model in advance, and no longer need a large number of cameras to capture the information such as eye rotation, facial expression and body change of the interactive object, and no longer need complex comprehensive data calculation to drive the three-dimensional character model. This low-cost, high-efficiency and more realistic remote interaction mode eliminates or weakens the obstacles of remote interaction and significantly improves the effect of remote interaction.
[0061] In fact, there are many similar solutions to the prior art introduced in the background section. For example, a mature technology is a remote interaction mode based on dynamic capture and holographic projection. This mode pre-stores shared virtual scenes and shared virtual models on the cloud, stores custom virtual scenes and virtual models in the user database, then uses a light motion capture device for motion capture, uses an audio input device for voice capture, and then establishes a logical host on the cloud for managing first-level logical slaves and second-level logical slaves to process the collected data, and finally uses three-dimensional holographic stereoscopic projection technology on the client to project the received cloud-processed stereoscopic scene holographic information onto the air or water mist, thereby realizing the interaction between the user and the three-dimensional image. This holographic three-dimensional stereoscopic projection technology is a 3D imaging technology that records and reproduces object light waves using laser interference and diffraction principles. The goal is to record all information of object light waves and present 3D virtual images in space through light diffraction and refraction. As can be seen, this way is technically complex and difficult to implement, requiring a large number of professional equipment and a special network. Except for some more professional or important occasions, it is not convenient for universal promotion. Compared with the simple, easy-to-implement, low-cost, and efficient remote interaction mode provided by the embodiments of the present application, it is obviously not advantageous.
[0062] For another example, there is a technology on the market that uses two cameras to capture real-time images to build a virtual three-dimensional model to realize remote restoration. This technology uses different cameras to capture the same object or person from different angles, then performs real-time calculation and rendering on the computer to generate a three-dimensional model of the object or person, and then the three-dimensional model is presented on the remote side through AR glasses, VR glasses, holographic projection, or naked-eye 3D display screen, etc. three-dimensional display devices. Although this mode has similarities with the embodiments of the present application, the principles are completely different. It still needs to rely on three-dimensional modeling and complex calculations, while the embodiments of the present application have broken away from these "formalities" and "light equipment" to achieve more realistic remote interaction.
[0063] The above describes the embodiments of the present application, but the above technical solutions can have many more optimized directions according to actual needs. For example, from the perspective of the position layout of the image capture device or the sound capture device, the "face-to-face" interaction effect can be further enhanced.
[0064] The first direction is described first, that is, considering the position of the image capture device. The layout between the first image capture device and the second image capture device mentioned above will affect the realism of the remote interaction between the two parties, and the horizontal distance between the first image capture device and the second image capture device is limited. In fact, the placement height of the first image capture device and the second image capture device and the vertical distance from the local side interaction object can also be specially set.
[0065] For the former, the reason why the height is considered is that in real face-to-face interaction, generally speaking, the two parties of the interaction mainly communicate face-to-face, that is, whether sitting or standing, the line of sight is basically on a horizontal plane, so that the mutual communication is normal, not awkward and sincere. Therefore, one feasible way to optimize the embodiment of the application is to set the first image capture device and the second image capture device within a predetermined pitch angle range of the horizontal line of sight plane in front of the interaction object on the side where the image capture device is located. The left and right eyes of the interaction object form two lines of sight, and the two lines of sight form a plane. If the image capture device is placed on the plane, when it captures the predetermined position of the local side interaction object, an angle of depression relative to the plane can be formed. If the image capture device is placed on the desktop, that is, on a plane lower than the aforementioned horizontal line of sight plane, when it captures the predetermined position of the local side interaction object, an angle of elevation relative to the desktop (i.e., a parallel plane of the aforementioned plane) can be formed. The predetermined pitch angle size here can be determined according to the normal pitch angle range of the line of sight of general personnel. See Figure 3 As shown, the image capture device is in front of the horizontal line of sight of the interaction object, and the range of the predetermined part of the interaction object (local side) captured by the image capture device is preferably within the normal angle of depression α or the range β of the angle of elevation, so that the interaction object (remote side) on the opposite side can see the interaction object (local side) through the image capture device, which is obviously natural and not easy to fatigue, and shows respect for the other party, thereby achieving effective simulation of real "face-to-face" interaction.
[0066] For the latter, the vertical distance between the first image capture device and the second image capture device relative to the local side interaction object (i.e., the distance from the midpoint of the line connecting the two image capture devices to the local side interaction object) cannot be too far, otherwise the simulated "face-to-face" communication will have too much "distance", nor can it be too close, otherwise the "face-to-face" communication will have too much "pressure", so it is necessary to choose a suitable distance. Therefore, the optimization direction of the embodiment of the application is to limit the vertical distance between the first image capture device and the second image capture device relative to the local side interaction object to be within the social distance range. The social distance depends on the way of social interaction. If it is a "one-on-one" interaction between two people, it can be kept within the social distance of two people, and if it is a large meeting with many people, the social distance can be the distance between the conference tables. See Figure 4 As shown, the social distance can be L1 or L2, L1 is less than L2, and as to how to choose within the interval, it can be determined by factors such as the way of interaction between the interaction objects, the degree of intimacy, the formality of the meeting, etc.
[0067] In addition to optimizing from the perspective of the image capture device, optimization can also be considered from another perspective, that is, the sound capture device (sound release device). As mentioned above, the sound capture device is used to capture the sound in the environment of the interaction object on the side where it is located (including the sound emitted by the interaction object itself, etc.). Although the sound capture device can be placed anywhere in the environment of the interaction object as long as it can effectively capture the sound, just like in face-to-face communication, sound cannot come from a place other than in front of the interaction object (the direction of the line of sight of the interaction object), and thus, in order to more realistically simulate face-to-face communication, as shown in FIG. 15, embodiments of the present application can consider placing the sound release device within a certain range of the middle of the lateral spacing between the two image capture devices, and generally, it should be below the middle, that is, simulate the position of the mouth below the eyes of the opposite interaction object. The two image capture devices (the first image capture device and the second image capture device) and the sound release device actually outline the face of the opposite interaction object, thereby more realistically displaying the interaction of the two parties. For the same reason, when the sound capture device and the sound release device are a whole product, or even different product components, the sound capture device can also consider the same optimization direction as described above. Figure 5
[0068] In addition, embodiments of the present application can also consider various situations encountered in reality and then optimize based on the needs of these situations themselves. For example, between the two objects interacting, it is generally not possible to remain completely still, and during the interaction, one or both of them will have body movements such as turning their heads, standing up and walking, etc. Especially for the interaction object on the side of the first image release device and the second image release device, since the images released by the image release devices are projected in front of the left and right eyes, in these cases, the pictures captured by the first image capture device and the second image capture device can follow the head of the user and will bring the understanding that the position of the "person on the other side" is drifting with the head of the user, thereby causing a sense of unreality.
[0069] In order to more clearly illustrate embodiments of the present application, further description will be made in conjunction with actual scenarios. Referring to FIG. 16, Figure 6a In this figure, a scenario of interaction between person A and person B is shown. Camera 1 (camera 3) and camera 2 (camera 4) are arranged in front of the line of sight of person A (person B) to capture images in the range of the scene where person A (person B) is located, including the face and upper body of person A (person B) itself, and loudspeaker 1 (or other device that can have both sound collection or sound emission) (loudspeaker 2) is arranged at a lower position in the middle of the lateral spacing between camera 1 (camera 3) and camera 2 (camera 4) to release the sound emitted by person B (person A). In this scenario, the related devices are arranged according to the following requirements:
[0070] (1) In the figure, two cameras correspond to the two eyes of the person B, and the horizontal distance between them is determined by the interpupillary distance of the person B. Of course, the interpupillary distance (i.e., the distance between the two pupils) of different people in reality is different, but it should not be too different, so as long as the distance between the two cameras corresponds to the interpupillary distance of the person B. For example, if the interpupillary distance of a man is about 6.1 cm and the interpupillary distance of a woman is about 5.8 cm, the horizontal distance between the two cameras can be set to a reference distance of 6 cm, and then adjusted within a range of ±3 cm using an adjustment knob to meet the needs of different people or even non-human groups.
[0071] (2) In the figure, the position of the loudspeaker 1 is also not randomly placed. Especially for devices that have both sound receiving and sound emitting functions, it is placed below the middle position of the horizontal distance between the two cameras, which corresponds to the mouth of the person B. The specific position can be determined by the distance between the midpoint of the horizontal distance between the eyes of the person B and the mouth.
[0072] (3) In the figure, the vertical distance between the two cameras and the person A cannot be too far or too close. To simulate face-to-face social interaction, the distance should be sufficient to meet the social distance. Of course, the social distance depends on the interaction method and closeness of the interacting parties. In this scenario, the person A and the person B are engaged in one-on-one communication, and the social distance can be set to about 1 m. If in other situations, the social distance can also be considered to be within the range of 0.2 m to 5 m.
[0073] (4) In the figure, the person A is sitting on a chair, and there is an office desk in front of him. The two cameras, as well as the microphone, loudspeaker, and other devices, are placed on the desktop. In fact, as mentioned earlier, as long as the placement of the cameras meets the normal human eye angle, the cameras placed on the desktop may not be directly on the same horizontal line as the person A's line of sight, but the person A's line of sight will not be too high above the desktop, so it can also meet the needs of interaction.
[0074] Since the scenario is a symmetrical interaction scenario, the layout of each device at the location of person B can be the same as that at the location of person A, and will not be described again. After completing the layout of each device described above, the camera 1 at the location of person A corresponds to the left eye of person B, and the camera 2 corresponds to the right eye of person B. The images recorded by the two cameras in real time are transmitted through the network, and then, through the corresponding three-dimensional display device (not shown in the figure), the images are projected to the left eye and the right eye of person B at the location of person B. The three-dimensional display device here can have many different implementation manners: the first is to use AR or VR glasses with different images for the left and right eyes, to project different images to the left and right eyes of person B, so that person B sees a three-dimensional image; the second is to use a naked-eye 3D display screen, to project different images for the left and right eyes to the left and right eyes of person B through the naked-eye 3D device, to achieve the effect of seeing a 3D image; and the third is to use a grating perspective 3D glasses to achieve the separate projection of images for the left and right eyes.
[0075] Through the foregoing device layout, when person B sees the image of person A, it is similar to the case where person B stands at the positions of the camera 1 and the camera 2 and sees the image, which is real and three-dimensional. In addition, when person A looks at the positions of the camera 1 and the camera 2 with the eyes, person B feels that person A is looking at B, so that person A and B can communicate through eye contact, to achieve a better remote interaction effect. Figure 6b The simulation effect generated by the foregoing interaction manner is shown, and person A and person B very realistically simulate the case of “face-to-face” communication.
[0076] The embodiments of the present application can use the foregoing hardware to achieve all the purposes of the invention, and therefore mainly introduce the hardware part. Those skilled in the art should understand that the embodiments of the present application can be provided as devices, systems, or related computer program products. Therefore, the present application can be implemented in a completely hardware embodiment, or in a completely software embodiment, or in an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.
[0077] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0078] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks.
[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0080] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0081] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, for storing, in general, data and / or program instructions. The memory can also include non-volatile memory, such as read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or non-volatile random access memory (NVRAM) for storing, in general, data and / or program instructions. The memory is an example of computer readable media.
[0082] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0083] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.
[0084] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.
Claims
1. A multi-captor based remote virtual interaction device, characterized in that, The device comprises: The first type of capture includes at least one group of image capture, each group of image capture includes a first image capture, a second image capture, for capturing image information including the predetermined position of the interaction object on the side of the first type of capture, the lateral distance between the first image capture and the second image capture is a first predetermined distance, the first predetermined distance is within the eye distance range of the interaction object on the side of the first type of capture; The second type of capture includes at least one sound capture, for capturing sound information including the position of the interaction object on the side of the second type of capture; The first type of release includes at least one group of image release, each group of image release includes a first image release, a second image release, for presenting the image information captured by the first image capture and the second image capture on the side of the first type of release respectively within the left and right eye view range of the interaction object on the side of the first type of release; The second type of release includes at least one sound release, for releasing the sound information captured by the second type of capture on the side of the second type of release to the interaction object on the side of the second type of release.
2. The apparatus of claim 1, wherein, The first image capture and the second image capture are arranged within a predetermined pitch angle range of the horizontal view plane in front of the interaction object on the side of the first type of capture.
3. The apparatus of claim 1, wherein, The vertical distance between the first image capture and the second image capture relative to the interaction object on the side of the first type of capture is a second predetermined distance, and the second predetermined distance is within the social distance range of the interaction object in real interaction.
4. The apparatus of claim 1, wherein, The sound release is arranged below the middle part of the lateral distance between the first image capture and the second image capture by a third predetermined distance, and the third predetermined distance is within the distance range between the eyes and the mouth of the interaction object on the side of the first type of capture.
5. A multi-captor based remote virtual interaction device, characterized in that, The device comprises: The first type of capture includes at least one group of image capture, each group of image capture includes a first image capture, a second image capture, for capturing image information including the predetermined position of the interaction object on the side of the first type of capture, the lateral distance between the first image capture and the second image capture is a first predetermined distance, the first predetermined distance is within the eye distance range of the interaction object on the side of the first type of capture; The second type of capture includes at least one sound capture, for capturing sound information including the position of the interaction object on the side of the second type of capture, so as to be transmitted to the corresponding sound release in the second type of release on the side of the second type of capture through the network transmission device.
6. A multi-captor based remote virtual interaction device, characterized in that, The device comprises: The first type of release includes at least one set of image release, each set of image release includes a first image release and a second image release, for presenting the image information of the interactive object captured by the first image capture and the second image capture on the opposite side of the first type of release respectively within the visual range of the left and right eyes of the interactive object on the side of the first type of release; The second type of release includes at least one sound release, for releasing the sound information captured by the second type of capture to the interactive object on the side of the second type of release.
7. The apparatus of any one of claims 1 to 6, wherein, The first image capture and the second image capture are cameras, and / or the sound capture in the second type of capture is a microphone, the sound release in the second type of release is a loudspeaker, and / or the image release in the first type of release is a three-dimensional display device.
8. The apparatus of claim 7, wherein, The three-dimensional display device is AR / VR glasses, and / or naked-eye 3D display screen, and / or grating 3D perspective glasses.
9. A multi-captor based remote virtual interaction system, characterized in that, The system includes: The first type of capture on the side of the remote interaction includes at least one set of image capture, each set of image capture includes a first image capture and a second image capture, for capturing image information including the predetermined position of the interactive object on the side of the first type of capture, the lateral distance between the first image capture and the second image capture is a first predetermined distance, and the first predetermined distance is within the interpupillary distance range of the interactive object on the opposite side of the first type of capture; The second type of capture on the side of the remote interaction includes at least one sound capture, for capturing sound information at the position of the interactive object on the side of the second type of capture; The first type of release on the other side of the remote interaction includes at least one set of image release, each set of image release includes a first image release and a second image release, for presenting the image information captured by the first image capture and the second image capture on the opposite side of the first type of release respectively within the visual range of the left and right eyes of the interactive object on the side of the first type of release; The second type of release on the other side of the remote interaction includes at least one sound release, for releasing the sound information captured by the second type of capture to the interactive object on the side of the second type of release; The network transmission device is used for transmitting the image information at the predetermined position of the interactive object captured by the first type of capture, and the sound information at the position of the interactive object captured by the second type of capture to the first type of release and the second type of release.
10. A multi-captor based remote virtual interaction system, characterized in that, The system includes: The first type of capture includes at least one set of image capture, each set of image capture includes a first image capture and a second image capture, for capturing image information including the predetermined position of the interactive object on the side of the first type of capture, the lateral distance between the first image capture and the second image capture is a first predetermined distance, and the first predetermined distance is within the interpupillary distance range of the interactive object on the opposite side of the first type of capture; The second type of capturer includes at least one sound capturer, which is configured to capture sound information at a position of an interactive object on a side where the second type of capturer is located; The first type of releaser includes at least one group of image releasers, each group of image releasers including a first image releaser and a second image releaser, which are configured to present image information captured by a first image capturer and a second image capturer on a side opposite to a side where the first type of releaser is located, respectively, to a left eye and a right eye of an interactive object on the side where the first type of releaser is located; The second type of releaser includes at least one sound releaser, which is configured to release sound information captured by the second type of capturer on a side opposite to a side where the second type of releaser is located, to an interactive object on the side where the second type of releaser is located; The network transmission device is configured to transmit image information captured by the first type of capturer at a predetermined position of an interactive object on a side where the first type of capturer is located, and sound information captured by the second type of capturer, to the first type of releaser and the second type of releaser on a side opposite to a side where the first type of capturer is located.
11. The system of claim 9 or 10, wherein, The first image capturer and the second image capturer are cameras, and / or the sound capturer in the second type of capturer is a microphone, the sound releaser in the second type of releaser is a loudspeaker, and / or the image releaser in the first type of releaser is a three-dimensional display device, the three-dimensional display device is AR / VR glasses, and / or a naked-eye 3D display screen, and / or a grating 3D perspective glasses.