Information processing apparatus and method, and computer readable storage medium

CN119998766APending Publication Date: 2025-05-13SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071040.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-10
Filing Date
2023-10-07
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the current post-production of films, special effects such as wand waving require post-production synthesis, which results in high time and money costs and is not conducive to the iteration of the shooting process.

Method used

By using processing circuits and multi-sensors in information processing equipment to achieve real-time interaction between real space and virtual space, special effects can be generated instantly during shooting and avoid post-production synthesis.

Benefits of technology

It greatly reduces the time and cost of film post-production, improves the iteration efficiency of the shooting process, and realizes real-time interactive virtual production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998766A_ABST
    Figure CN119998766A_ABST
Patent Text Reader

Abstract

The invention relates to an information processing apparatus and method, and a computer readable storage medium. The information processing apparatus includes processing circuitry configured to, in response to a trigger based on at least one element in a real space, generate a trigger event in a virtual space corresponding to the real space, thereby enabling real-time interaction between the real space and the virtual space.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and method, and computer-readable storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on October 10, 2022, with application number 202211233436.6 and invention name “Information processing device and method, computer-readable storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of information processing technology, and more particularly to achieving real-time interaction between a real space and a virtual space corresponding to the real space, and more particularly to an information processing device and method, and a computer-readable storage medium. Background Art

[0003] Although virtual production technology has greatly reduced the requirements for video post-production, some special effects still require post-production. For example, in film post-production, if an actor waves a magic wand that produces special effects, the position of the magic wand cannot be strictly preset during filming. Therefore, existing technology can only achieve this through post-production synthesis of special effects. In other words, film post-production usually requires a high cost in time and money, and the special effects cannot be seen until after synthesis, which is very unfavorable for the iteration of the filming process.

[0004] Summary of the Invention

[0005] A brief overview of the present invention is provided below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify key or important aspects of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is simply to present certain concepts in a simplified form as a prelude to the more detailed description discussed later.

[0006] According to one aspect of the present disclosure, an information processing device is provided, which includes a processing circuit configured to: generate a trigger event in a virtual space corresponding to the real space in response to a trigger based on at least one element in the real space, thereby realizing real-time interaction between the real space and the virtual space.

[0007] In the information processing device according to the embodiment of the present disclosure, when a trigger is performed based on at least one element in the real space, a trigger event will be generated accordingly in the virtual space, thereby enabling real-time interaction between the real space and the virtual space, and further enabling real-time interactive virtual production.

[0008] According to another aspect of the present disclosure, an information processing method is provided, comprising: generating a trigger event in a virtual space corresponding to the real space in response to a trigger based on at least one element in the real space, thereby realizing real-time interaction between the real space and the virtual space.

[0009] According to other aspects of the present invention, there are also provided computer program codes and computer program products for implementing the above-mentioned information processing method, as well as a computer-readable storage medium having recorded thereon the computer program codes for implementing the above-mentioned information processing method. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to further illustrate the above and other advantages and features of the present invention, the following is a further detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings. The accompanying drawings, together with the detailed description below, are included in this specification and form a part of this specification. Elements with the same function and structure are represented by the same reference numerals. It should be understood that these drawings only depict typical examples of the present invention and should not be regarded as limiting the scope of the present invention. In the drawings:

[0011] FIG1 shows a functional module block diagram of an information processing device according to an embodiment of the present disclosure.

[0012] 2A-2F are exemplary diagrams illustrating virtual film and television production using an information processing device according to an embodiment of the present disclosure.

[0013] 3A-3D are exemplary diagrams illustrating virtual production of an online class using an information processing device according to an embodiment of the present disclosure.

[0014] FIG4 is a flowchart illustrating an example of the flow of an information processing method according to an embodiment of the present disclosure.

[0015] FIG. 5 is a block diagram showing an example structure of a personal computer that can be employed in the embodiments of the present disclosure. DETAILED DESCRIPTION

[0016] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings. For the sake of clarity and conciseness, not all features of an actual implementation are described in this specification. However, it should be understood that in the process of developing any such actual implementation, many implementation-specific decisions must be made in order to achieve the developer's specific goals, such as compliance with system and business-related constraints, which may vary from implementation to implementation. In addition, it should be understood that although the development work may be very complex and time-consuming, it is a routine task for those skilled in the art who benefit from the contents of this disclosure.

[0017] It is also necessary to explain here that, in order to avoid obscuring the present disclosure due to unnecessary details, the accompanying drawings only show the device structure and / or processing steps closely related to the scheme according to the present disclosure, while other details that are not closely related to the present disclosure are omitted.

[0018] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0019] Figure 1 shows a functional module block diagram of an information processing device 100 according to an embodiment of the present disclosure. As shown in Figure 1, the information processing device 100 includes: a processing unit 102, which can be configured to generate a trigger event in a virtual space corresponding to the real space in response to a trigger based on at least one element in the real space, thereby realizing real-time interaction between the real space and the virtual space.

[0020] The processing unit 102 may be implemented by one or more processing circuits, which may be implemented as a chip, for example.

[0021] As an example, the real space may also be referred to as a real world or a real scene, and the virtual space may also be referred to as a virtual world or a virtual scene.

[0022] The at least one element includes one or more of an entity in a real space (eg, a person, an animal, another object, etc.), a background, light, and sound.

[0023] For example, the real space can be the actual filming space in a film or television production, and the virtual space can be a predetermined virtual set in the film or television production. For example, actors can interact with objects, environments, lighting, and other elements in the virtual set through body movements, gestures, and expressions; actors can interact with objects, environments, lighting, and other elements in the virtual set through props; actors can interact with objects, environments, lighting, and other elements in the virtual set through specific sounds; and actors' voices can be changed or transformed in real time to achieve interaction.

[0024] As an example, the real space can be a real shooting space in an online class, and the virtual space can be a predetermined virtual scene in the online class. For example, the teacher can interact with objects in the virtual scene by using a pointer.

[0025] Those skilled in the art can also think of other video scenarios besides film and television production and online classes, which will not be repeated here.

[0026] Take the example of a magic wand-waving special effect in film and television production. Because it's impossible to predict the actor's wand-waving position and posture, existing methods require filming the actor's wand-waving motion and then overlaying the wand effect on the film in post-production. This post-production process is often time-consuming and expensive, and the special effect can only be seen after compositing, which greatly hinders the iteration process of the filming process.

[0027] In the information processing device 100 according to an embodiment of the present disclosure, when a trigger is performed based on at least one element in the real space, a trigger event will be generated accordingly in the virtual space, thereby enabling real-time interaction between the real space and the virtual space, and further enabling real-time interactive virtual production.

[0028] Taking the above-mentioned magic wand waving special effects in film and television production as an example, the information processing device 100 according to the embodiment of the present disclosure enables real-time interaction between the real space and the virtual space, so that when the actor waves the magic wand in the real space, real-time interaction can be generated in the virtual space, that is, real-time interactive virtual production can be realized, thereby largely avoiding the work of post-production synthesis, thereby saving time and money costs.

[0029] In the following, for convenience, the description is made using a film and television production scenario as an example.

[0030] As an example, the processing unit 102 may be configured to simulate a predetermined area in the real space that includes at least one element to construct a simulated area, and to enable real-time interaction based on the association between the simulated area and the virtual space. Thus, the predetermined area is simulated, and the simulated area serves as a bridge between the predetermined area in the real space and the virtual space.

[0031] For example, in film and television production, the predetermined area may be a shooting area, and simulating the predetermined area refers to simulating or emulating the shooting area (for example, digitizing the state of the shooting area) to construct a simulation area (for example, a digitized area) corresponding to the shooting area.

[0032] For example, the processing unit 102 may map (or convert) objects in the simulation area into the virtual space, thereby establishing an association between the simulation area and the virtual space.

[0033] For example, the processing unit 102 can obtain the positions of all objects in the simulation area in the real-world coordinate system in real time through the following multi-sensor joint calibration, so as to establish a connection between the real space and the virtual space through the simulation area, thereby realizing real-time interaction between the real space and the virtual space.

[0034] As an example, the processing unit 102 may be configured to perform simulation processing based on perception data obtained by at least one sensor (eg, a data sensor) sensing a predetermined area.

[0035] For example, at least one sensor collects data from a predetermined area in real time to obtain perception data. For example, in film and television production, multiple sensors are deployed in the shooting area to perceive changes in the scene. The perception forms include sound, light, video, audio, distance, motion, etc.

[0036] The processing unit 102 can process and analyze the sensory data in real time to perform simulation processing. For example, in film and television production, the processing unit 102 simulates actors, props, and scenes in real space based on the sensory data (for example, by analyzing the coordinates, posture, movement, expression, and voice of actors, props, and other elements in real time), thereby constructing a simulation area.

[0037] As an example, the at least one sensor includes one or more of a visible light sensor, an infrared sensor, an ultraviolet sensor, a distance sensing sensor, an audio sensor, and a motion sensor.

[0038] For example, visible light sensors, infrared sensors, and ultraviolet sensors are video sensors. They can be used to capture video or image information of a predetermined area in real space, providing raw data for video signal processing. Visible light sensors, infrared sensors, and ultraviolet sensors can be included in cameras, video cameras, video capture cards, and the like. The wavelengths they capture include, but are not limited to, visible light, infrared, and ultraviolet.

[0039] For example, an audio sensor is used to collect audio information within a predetermined area in a real space and provide raw data for audio signal processing. The number of audio sensors is at least 1. The audio sensor may be, for example, a microphone.

[0040] For example, distance perception sensors are used to perceive the distance from all objects (actors, props, scenery, etc.) in a predetermined area to each sensor in real time; distance perception sensors can be combined with point cloud recognition and other algorithms to enhance the accuracy of identifying, positioning and tracking objects and human bodies; distance perception sensors can be combined with video signal processing and multi-sensor joint calibration to obtain the positions of all objects in the predetermined area in the real-world coordinate system in real time, thereby establishing an association between the real-world coordinate system and the virtual world coordinate system through the simulation area (that is, establishing an association between the real space and the virtual space) so that the coordinates of objects in the predetermined area can be calculated more accurately; distance sensors include but are not limited to: LIDAR, I-TOF sensor, D-TOF sensor, structured light distance sensor, millimeter wave distance sensor, etc.

[0041] For example, a motion sensor is used to sense the motion of all objects within a predetermined area.

[0042] The above-mentioned sensor can sense a predetermined area in the real space in an omnidirectional manner.

[0043] As an example, the processing unit 102 may be configured to analyze at least one element included in the simulation area and trigger when the at least one element meets a predetermined trigger condition, wherein the at least one element included in the simulation area corresponds to the at least one element included in the real space. Thus, by analyzing the elements included in the simulation area, different state changes of elements in a predetermined area in the real space are analyzed, and interaction between the real space and the virtual space is triggered when the predetermined trigger condition is met.

[0044] For example, the predetermined trigger condition may be that an actor waves a magic wand, the light change in the predetermined area exceeds a predetermined threshold, and so on.

[0045] As an example, the at least one element includes one or more of entities, background, light, and sound in the simulated area. For example, the entities in the simulated area may correspond to people, animals, other objects, etc. in a predetermined area included in the real space; the background in the simulated area may correspond to the background in the predetermined area; the light in the simulated area may correspond to the light in the predetermined area; and the sound in the simulated area may correspond to the sound in the predetermined area.

[0046] The processing unit 102 can analyze changes in all elements in the simulation area in real time (for example, analyzing actors, props, background, lighting, and sounds captured in the scene), track and capture key elements, and obtain the spatial position, posture, and motion state of key elements in real time (for actors, facial recognition and analysis of the actor's facial expressions can also be performed). The state of the target object is compared with a pre-determined trigger condition. When the target object state meets the pre-determined trigger condition, real-time interaction with the virtual scene is triggered (i.e., a trigger is performed).

[0047] As an example, the processing unit 102 may be configured to determine whether the at least one element satisfies a predetermined trigger condition based on element information related to the at least one element, thereby performing a trigger.

[0048] As an example, the element information may include one or more of video information, audio information, and three-dimensional information related to at least one element.

[0049] As an example, the video information may include one or more of face recognition information, object recognition information, gesture information, expression information, and posture information.

[0050] For example, the processing unit 102 can perform face recognition, positioning and tracking, such as being used to distinguish, locate and track the identity IDs of the interactive actors, and assign different interactive content to different IDs. For example, actors A and B wave their wands at the same time, but the special effects of A and B's wands are different. For example, when the processing unit 102 is used to perform face recognition, positioning and tracking, it can include but is not limited to the following functions. Face entry function: entering the faces of actors who are not in the database to ensure that all actors can be accurately identified; face recognition function: identifying all people appearing in the image and giving their identity IDs; face positioning function: locating the faces appearing in the image and outputting the corresponding coordinates of the faces in the image; face tracking function: tracking the faces moving in the image and dynamically outputting accurate face coordinates.

[0051] Face recognition, positioning and tracking can be achieved through various appropriate technologies. For example, please refer to the technology described in the document "Sefik Ilkin Serengil et al., LightFace: A Hybrid Deep Face Recognition Framework, 2020 Innovations in Intelligent Systems and Applications Conference (ASYU), 2020", which will not be repeated here.

[0052] For example, the processing unit 102 can perform target object recognition, positioning and tracking, such as for identifying special objects in a video signal, and positioning and tracking the objects, such as to facilitate rendering real-time special effects for these objects (such as a wand held by an actor, a glowing torch, a prop gun emitting lasers, etc.). For example, when the processing unit 102 is used to perform target object recognition, positioning and tracking, it can include but is not limited to the following functions. Object entry function: entering objects that are not in the database to ensure that all special objects can be accurately identified; object recognition function: identifying special objects appearing in the image; object positioning function: locating objects appearing in the image and outputting the corresponding coordinates of the object in the image; object tracking function: tracking the face of an object moving in the image and dynamically outputting accurate object coordinates.

[0053] Target object recognition, positioning, and tracking can be achieved through various appropriate technologies, for example, refer to the literature "Alexey Bochkovskiy et al., YOLOv4: Optimal Speed ​​and Accuracy of Object Detection, arXiv:2004.10934, 2020" and the technology described in https: / / github.com / ultralytics / yolov5, which will not be repeated here.

[0054] For example, the processing unit 102 can perform human key point detection, such as for identifying a human body in a video signal and extracting the key points of the human body appearing in the video signal, so as to facilitate triggering some interactive effects through certain specific actions of the actor (for example, triggering an explosion effect when the actor squats, etc.). For example, when the processing unit 102 is used to perform human key point detection, it can include but is not limited to the following functions. Human key point detection function: extracting key skeletal coordinate points of the human body appearing in the video signal; human behavior analysis function: analyzing the behavior of the human body in the video signal, such as walking, squatting, running, etc.

[0055] The detection of key points on the human body can be achieved through various appropriate technologies, such as those described in https: / / developer.huawei.com / consumer / en / doc / development / hiai-Guides / skel eton-detection-0000001051008415, which will not be repeated here.

[0056] For example, the processing unit 102 can perform gesture recognition, such as for identifying and tracking hands in video signals, identifying gesture information corresponding to the hands, so as to trigger some interactive special effects by having actors complete specific gestures (for example, the actor snaps his fingers to trigger ambient light change effects, etc.). For example, when performing gesture recognition, the processing unit 102 may include but is not limited to the following functions. Gesture entry function: entering gestures that are not in the database to ensure that all gestures can be accurately recognized; hand recognition function: identifying hands appearing in the image; hand positioning function: locating hands appearing in the image, and outputting the corresponding coordinates of the hands in the image; hand tracking function: tracking the moving hands in the image, and dynamically outputting accurate hand coordinates; gesture recognition function: recognizing gesture information expressed by the hands, such as making a heart, snapping fingers, and giving a thumbs-up.

[0057] Gesture recognition can be implemented through various appropriate technologies, such as those described in https: / / github.com / yeemachine / kalidokit, which will not be described here.

[0058] For example, the processing unit 102 can perform facial expression recognition, for example, to identify the facial expressions of a tracked actor, so as to trigger certain interactive special effects by having the actor perform a specific expression (e.g., rendering a flashing heart-shaped special effect on the screen when the actor expresses happiness). For example, when performing facial expression recognition, the processing unit 102 may include, but is not limited to, the following functions: Expression entry function: Entering expressions not in the database to ensure that all expressions can be accurately recognized; Expression recognition function: Recognizing facial expression information on the actor, such as happiness, sadness, anger, fear, etc.

[0059] Expression recognition can be achieved through various appropriate technologies, such as those described in https: / / github.com / yeemachine / kalidokit, which will not be repeated here.

[0060] As an example, the audio information includes voiceprint recognition information and / or sound conversion information.

[0061] For example, the processing unit 102 can perform sound recognition, for example, to identify certain specific sounds in an audio signal, so as to facilitate triggering certain interactive special effects through these specific sounds (e.g., the sound produced when an actor strikes a wooden barrel filled with explosives triggers an explosion effect, etc.). For example, when performing sound recognition, the processing unit 102 may include, but is not limited to, the following functions: Sound recording function: recording sounds that are not in the database to ensure that all specific sounds can be accurately recognized; Sound recognition function: identifying specific sounds emitted by a predetermined area included in a real space and providing the accurate time when the sound was emitted.

[0062] For example, the processing unit 102 can perform speech recognition, for example, to identify specific lines spoken by an actor in an audio signal, thereby facilitating the triggering of certain interactive special effects through certain specific lines (e.g., an actor speaking a spell to trigger certain magical special effects, etc.). For example, when performing speech recognition, the processing unit 102 may include, but is not limited to, the following functions: Voice input function: to input lines that are not in the database to ensure that all specific lines can be accurately recognized; Voice recognition function: to identify specific lines spoken in the shooting area and provide the accurate time when the lines were spoken.

[0063] For example, processing unit 102 can perform voice conversion, for example, converting the voice of an actor in an audio signal in real time, such as performing voice-changing processing on the actor. For example, processing unit 102 can include, but is not limited to, the following functions when performing voice conversion: Voiceprint recognition function: distinguishing the voices of different actors so that only the voice of a specific actor can be processed; Voice conversion function: processing the actor's voice in real time to achieve different sound effects.

[0064] As an example, the three-dimensional information includes one or more of spatial position information, motion information, and distance information.

[0065] For example, processing unit 102 can perform motion capture, e.g., to capture an actor's movements in real time. Motion capture works in conjunction with human key point detection to achieve more accurate and faster actor motion capture. Motion capture is particularly complementary for difficult movements or visual blind spots. The motion capture function can be implemented using a specific motion capture module. For example, a motion capture module is typically miniaturized and can be discreetly worn on the actor, without interfering with or affecting the filming process. Motion capture modules include, but are not limited to, inertial device-based motion capture devices, video signal-based or depth sensor-based motion capture devices, and visible or invisible landmark-based motion capture devices.

[0066] For example, the processing unit 102 can perform multi-sensor joint calibration, for example, by photographing a target whose size or three-dimensional structure is completely known, to achieve joint calibration of video, audio, distance, motion and other sensors, obtain the position and posture of all sensors in the real-world coordinate system, and then obtain the accurate position of objects in a predetermined area included in the real space in the real-world coordinate system; and by adjusting the virtual world coordinate system, the virtual world coordinate system can be aligned with the real-world coordinate system.

[0067] The joint calibration of multiple sensors can be achieved through various appropriate technologies. For example, please refer to the document "Guohang Yan et al., OpenCalib: A Multi-sensor Calibration Toolbox for Autonomous Driving, arXiv: 2205.14087, 2022", which will not be repeated here.

[0068] For example, the element information can be used to provide metadata for the simulation process. For example, the processing unit 102 can use artificial intelligence and other technologies to analyze the status of the actors, props, and scenes in the predetermined area to provide metadata for the simulation process.

[0069] In addition, after constructing the simulation area, the processing unit 102 can analyze the changes in at least one of the coordinates of the actors, props, and scenes in the simulation area, or analyze at least one of the changes in the actors' movements, postures, expressions, gestures, etc., or analyze at least one of the changes in the actors' voices, scene sounds, etc., to determine whether the predetermined trigger conditions are met.

[0070] For example, the processing unit 102 may also process the above-mentioned element information and perform interactive actions with the constructed virtual world according to preset script rules.

[0071] As an example, the processing unit 102 may be configured to render the virtual space based on a triggering event generated by the real-time interaction, thereby obtaining an updated virtual space.

[0072] For example, a triggering event might be the generation of film or television special effects. Examples include object collisions, explosions, and sparks. For example, if actor A is recognized as waving a wand, special effects rendering at the wand tip in the virtual world is triggered. For example, if actor B is recognized as moving with a torch, a moving light source is added to the corresponding location in the virtual world. For example, if actor C is recognized as expressing sadness, the primary light source in the virtual world is dimmed, and dark clouds are added to create atmosphere. This allows for full, three-dimensional interaction with the virtual space.

[0073] As an example, the trigger event may also be an online teaching special effect, such as a special effect related to teacher-student interaction in an online classroom scenario. Those skilled in the art will also think of other examples of trigger events, which will not be repeated here.

[0074] As an example, the processing unit 102 may be configured to display the updated virtual space on a display device that displays the virtual space instead of the virtual space, thereby obtaining an updated display image related to the triggering event.

[0075] As an example, the display device includes an LED screen (light emitting diode screen). The LED screen can also be referred to as an LED VP (virtual production) screen or a VP large screen. Those skilled in the art can also conceive of other examples of display devices, which will not be repeated here.

[0076] For example, the updated display image can be used as a real-time shooting background for film and television shooting.

[0077] For example, during VP filming, the shooting area in real space is simulated (e.g., digitized) to enable real-time interaction between the shooting area and the virtual space. The rendered virtual space is then displayed on a large LED screen as the VP background, achieving real-time interactive VP special effects. In other words, the information processing device 100 uses real-to-virtual technology to enable interaction between the real shooting area and the virtual world, and renders the interaction results and special effects on the LED VP screen in real time.

[0078] As an example, the processing unit 102 may be configured to use a camera to capture a predetermined area when a trigger event occurs and an image displayed by a display device (eg, an image with special effects related to the trigger event) to obtain a captured image.

[0079] For example, the processing unit 102 may use an image obtained by photographing the predetermined area when the trigger event occurs as the foreground image, and an image obtained by photographing the image displayed by the display device as the background image. By generating a trigger event based on changes in the shooting scene (e.g., changes in elements in the predetermined area) to update the image displayed by the display device (i.e., changing the shooting background via the updated display image), real-time interactive virtual shooting can be achieved.

[0080] As an example, the processing unit 102 may be configured to determine the position and / or size of the rendering in the virtual space based on the position and / or posture relationship between the camera device and the display device. For example, the position and / or size of the rendering for generating special effects in the virtual space may be determined based on the position and / or posture relationship between the camera device and the display device.

[0081] For example, for the scene mentioned above where the actor waves a magic wand that produces special effects, through real2virtual technology, the information processing device 100 captures the position and / or posture of the actor and the magic wand, and combines the position and / or posture relationship between the camera device and the display device to interact this information with the virtual space and produce special effects. Through real-time rendering of the virtual space, the image with the magic wand special effects accurately superimposed is displayed on the LED large screen, thereby completing the shooting of the magic wand special effects in real time, avoiding post-production, and saving time and money costs to a great extent.

[0082] For example, high-performance computers (clusters), high-performance graphics cards, cloud computing technology, etc. can be used to provide high-performance computing and processing capabilities for the information processing device 100.

[0083] 2A-2F are exemplary diagrams illustrating virtual film and television production using the information processing device 100 according to an embodiment of the present disclosure.

[0084] As shown in Figure 2A, actors 1 and 2 stand in front of an LED VR screen. The shooting area includes at least actors 1 and 2, and actor 1 holds a prop. Although not shown in Figure 2A, actor 1 can wear a motion capture device (e.g., MoCap) to capture their movements. Multiple sensors (e.g., at least one of a video sensor, an audio sensor, a distance sensing sensor, and a motion sensor) are used to collect data from the shooting area to obtain perception data, and the camera device captures the shooting area and the LED VR screen. As can be seen from the description above, the information processing device 100 can construct a simulation area corresponding to the shooting area based on the perception data.

[0085] As shown in Figure 2B, the information processing device 100 analyzes the simulation area, for example, identifies actor 1 and actor 2 respectively through face recognition, and identifies props (for example, the magic wand held by actor 1) through object recognition, and the information obtained about actor 1 and actor 2 is, for example: actor 1 holds a magic wand, stands with his left hand raised, has coordinates (600, 800) in the real space, and has a smiling expression; actor 2 holds a school bag, stands with his coordinates (300, 600) in the real space, and has a calm expression.

[0086] As shown in Figure 2C, the information processing device 100 performs audio analysis on the simulation area, and for example obtains information about two audios (for example, audio 1 and audio 2): the type of audio 1 is voice, and the voice recognition result is "Confringo", which is judged to be the voiceprint of actor 1; the type of audio 2 is sound, and the recognition result is ambient sound.

[0087] As shown in Figure 2D , information processing device 100 performs position and motion analysis on the simulation area. For example, the following information is obtained about actors 1 and 2: actor 1's position within the simulation area is (1.2, 22.5, 10.5), actor 2's position within the simulation area is (2.2, 22.3, 11.2), and the wand's position within the simulation area is (3.2, 25.0, 11.2); actor 1's motion is waving his left wrist. In Figure 2D , the position information is schematically represented by boxes.

[0088] As shown in FIG2E , a virtual space is pre-set and the preset virtual space (for example, the scene excluding actor 1, actor 2, and the camera device as shown in FIG2E ) is displayed on the LED VP screen, and the information processing device 100 associates the position information of the actors, props, and other elements in the simulation area with the virtual space, and establishes a bridge (association) between the shooting area and the virtual space through the simulation area.

[0089] As shown in Figure 2F, assume the preset trigger condition is "[Actor 1] holds [magic wand], waves his wrist, and receives the voice message [Confringo]." When this trigger condition is met, a trigger event is generated, rendering the virtual space. For example, the trigger event could be a movie or TV special effect: light particles rendered at the head of the magic wand, or objects exploding in the direction the wand is pointing. As mentioned above, the position and / or size of the rendering is determined based on the position and / or posture relationship between the camera and the LED VP screen.

[0090] 2A-2F , it can be seen that the information processing device 100 according to the embodiment of the present disclosure can realize real-time interaction between the shooting area and the virtual space.

[0091] 3A-3D are exemplary diagrams illustrating virtual production of an online class using the information processing device 100 according to an embodiment of the present disclosure.

[0092] FIG3A shows an example diagram summarizing the virtual production of an online classroom. As shown in FIG3A , the teacher teaches in the teaching area, and a plurality of sensors (for example, at least one of a video sensor, an audio sensor, a distance sensing sensor, and a motion sensor) are used to collect data from the teaching area to obtain perception data, and a camera device shoots the teaching area and the display screen. As can be seen from the description above, the information processing device 100 can construct a simulation area corresponding to the teaching area based on the perception data. The information processing device 100 can interact with the elements in the virtual space displayed on the display screen (for example, “current intensity”, “resistance”, and “Ohm’s law” shown in FIG3A ) in the classroom based on the automatic analysis of the teacher’s actions, behaviors, voice, and other information in the simulation area, and display the results of the interaction in real time on the student display terminal.

[0093] For example, as shown in Figure 3B, the information processing device 100 can analyze the position of the teacher's pointer in the simulation area and identify the teacher's actions such as tapping the screen and marking key points on the screen, and render the interaction results between the teacher and the virtual space displayed on the display screen in real time based on the identified actions, and display the interaction results in real time on the student display terminal.

[0094] For example, as shown in Figure 3C, the information processing device 100 can automatically analyze the blackboard writing written by the teacher through optical character analysis (OCR), perform real-time rendering of the content in the blackboard writing by superimposing special effects (for example, special effects that highlight the physical elements in Figure 3C), and display the rendering results in real time on the student display terminal.

[0095] For example, as shown in Figure 3D, the information processing device 100 can automatically analyze the voice of a teacher (e.g., Mr. Smith) and convert the teacher's voice (e.g., "Hi @John, do you know the answer?") and the translation results into subtitles ("Hi, @John, do you know the answer?") in real time, which are then displayed on the student's display terminal. Incorporating technologies such as natural language processing (NLP) allows for simple interaction with students.

[0096] 3A-3D , it can be seen that the information processing device 100 according to the embodiment of the present disclosure can realize real-time interaction between the teaching area and the virtual space.

[0097] Corresponding to the above-mentioned information processing device embodiments, the present disclosure also provides embodiments of information processing methods.

[0098] FIG. 4 is a flowchart illustrating an example of a process of an information processing method S400 according to an embodiment of the present disclosure.

[0099] The information processing method S400 according to an embodiment of the present disclosure starts from S402 .

[0100] In S404 , in response to a trigger based on at least one element in the real space, a trigger event is generated in the virtual space corresponding to the real space, thereby achieving real-time interaction between the real space and the virtual space.

[0101] The information processing method S400 ends at S406.

[0102] In the information processing method S400 according to an embodiment of the present disclosure, when a trigger is performed based on at least one element in the real space, a trigger event will be generated accordingly in the virtual space, thereby enabling real-time interaction between the real space and the virtual space, and further enabling real-time interactive virtual production.

[0103] As an example, in S404 , a predetermined area including at least one element in the real space is simulated to construct a simulation area, and real-time interaction is achieved based on the association between the simulation area and the virtual space.

[0104] As an example, simulation processing is performed based on perception data obtained by at least one sensor sensing a predetermined area.

[0105] As an example, the at least one sensor includes one or more of a visible light sensor, an infrared sensor, an ultraviolet sensor, a distance sensing sensor, an audio sensor, and a motion sensor.

[0106] As an example, at least one element included in the simulation area is analyzed, and a trigger is performed when the at least one element meets a predetermined trigger condition, and wherein the at least one element included in the simulation area corresponds to the at least one element included in the real space.

[0107] As an example, the at least one element includes one or more of an entity, a background, a light, and a sound in the simulation area.

[0108] As an example, based on element information related to at least one element, it is determined whether the at least one element meets a predetermined trigger condition, thereby triggering.

[0109] As an example, the element information includes one or more of video information, audio information, and three-dimensional information related to at least one element.

[0110] As an example, the video information includes one or more of face recognition information, object recognition information, gesture information, expression information, and posture information.

[0111] As an example, the audio information includes voiceprint recognition information and / or sound conversion information.

[0112] As an example, the three-dimensional information includes one or more of spatial position information, motion information, and distance information.

[0113] As an example, the virtual space is rendered based on a trigger event generated by the real-time interaction, thereby obtaining an updated virtual space.

[0114] As an example, the triggering event is a movie or TV special effect.

[0115] As an example, the updated virtual space is displayed on a display device that displays the virtual space instead of the virtual space, thereby obtaining an updated display image related to the triggering event.

[0116] As an example, a camera device is used to capture a predetermined area when a trigger event occurs and an image displayed by a display device, thereby obtaining a captured image.

[0117] As an example, the position and / or size of the rendering in the virtual space is determined based on the position and / or posture relationship between the camera device and the display device.

[0118] The information processing method S400 according to an embodiment of the present disclosure may be executed by the information processing device 100 described above. For specific details, please refer to the description of the related processing of the information processing device 100, which will not be repeated here.

[0119] The basic principles of the present invention are described above in conjunction with specific embodiments. However, it should be pointed out that those skilled in the art will understand that all or any steps or components of the methods and devices of the present invention can be implemented in any computing device (including a processor, storage medium, etc.) or a network of computing devices in the form of hardware, firmware, software, or a combination thereof. This can be achieved by those skilled in the art using their basic circuit design knowledge or basic programming skills after reading the description of the present invention.

[0120] Furthermore, the present invention also provides a program product storing machine-readable instruction codes. When the instruction codes are read and executed by a machine, the method according to the embodiment of the present invention can be executed.

[0121] Accordingly, the storage medium for carrying the program product storing the machine-readable instruction code is also included in the disclosure of the present invention. The storage medium includes but is not limited to a floppy disk, an optical disk, a magneto-optical disk, a memory card, a memory stick, and the like.

[0122] When the present invention is implemented through software or firmware, the programs constituting the software are installed from a storage medium or a network to a computer with a dedicated hardware structure (such as the general-purpose computer 500 shown in Figure 5). When various programs are installed on the computer, it can perform various functions, etc.

[0123] 5 , a central processing unit (CPU) 501 executes various processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage section 508 to a random access memory (RAM) 503. In the RAM 503, data required when the CPU 501 executes various processes, etc., is also stored as needed. The CPU 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output interface 505 is also connected to the bus 504.

[0124] The following components are connected to the input / output interface 505: an input section 506 (including a keyboard, a mouse, etc.), an output section 507 (including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.), a storage section 508 (including a hard disk, etc.), and a communication section 509 (including a network interface card such as a LAN card, a modem, etc.). The communication section 509 performs communication processing via a network such as the Internet. A drive 55 may also be connected to the input / output interface 505 as needed. A removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed in the drive 55 as needed, so that a computer program read therefrom is installed in the storage section 508 as needed.

[0125] In the case of realizing the above-described series of processing by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 511 .

[0126] It should be understood by those skilled in the art that such storage media is not limited to the removable medium 511 shown in FIG5 , which stores the program therein and is distributed separately from the device to provide the program to the user. Examples of the removable medium 511 include magnetic disks (including floppy disks (registered trademark)), optical disks (including compact disk read-only memories (CD-ROMs) and digital versatile disks (DVDs)), magneto-optical disks (including minidiscs (MDs) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be a ROM 502, a hard disk included in the storage section 508, or the like, in which the program is stored and distributed to the user together with the device containing them.

[0127] It should also be noted that in the apparatus, method, and system of the present invention, each component or step can be decomposed and / or recombined. Such decomposition and / or recombination should be considered equivalent solutions of the present invention. Furthermore, the steps of performing the above series of processes can naturally be performed in chronological order according to the order described, but do not necessarily need to be performed in chronological order. Certain steps can be performed in parallel or independently of each other.

[0128] Finally, it should be noted that the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. Furthermore, in the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0129] Although the embodiments of the present invention have been described in detail above with reference to the accompanying drawings, it should be understood that the embodiments described above are merely illustrative of the present invention and are not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations may be made to the embodiments described above without departing from the spirit and scope of the present invention. Therefore, the scope of the present invention is limited solely by the appended claims and their equivalents.

[0130] The present technology can also be implemented as follows.

[0131] Note 1. An information processing device comprising:

[0132] The processing circuit is configured to:

[0133] In response to a trigger based on at least one element in the real space, a trigger event is generated in the virtual space corresponding to the real space, thereby achieving real-time interaction between the real space and the virtual space.

[0134] Note 2. An information processing device according to Note 1, wherein the processing circuit is configured to simulate a predetermined area in the real space including the at least one element to construct a simulation area, and realize the real-time interaction based on the association between the simulation area and the virtual space.

[0135] Supplementary note 3. The information processing device according to Supplementary note 2, wherein the processing circuit is configured to perform the simulation processing based on perception data obtained by at least one sensor sensing the predetermined area.

[0136] Note 4. The information processing device according to Note 3, wherein the at least one sensor includes one or more of a visible light sensor, an infrared sensor, an ultraviolet sensor, a distance perception sensor, an audio sensor, and a motion sensor.

[0137] Note 5. An information processing device according to any one of Notes 2 to 4, wherein the processing circuit is configured to analyze at least one element included in the simulation area, and to perform the triggering when the at least one element meets a predetermined trigger condition, and wherein the at least one element included in the simulation area corresponds to the at least one element included in the real space.

[0138] Supplementary note 6. The information processing device according to Supplementary note 5, wherein the at least one element includes one or more of an entity, a background, light, and sound in the simulation area.

[0139] Note 7. The information processing device according to Note 5 or 6, wherein the processing circuit is configured to determine whether the at least one element meets the predetermined trigger condition based on element information related to the at least one element, thereby performing the triggering.

[0140] Supplementary note 8. The information processing apparatus according to Supplementary note 7, wherein the element information includes one or more of video information, audio information, and three-dimensional information related to the at least one element.

[0141] Note 9. The information processing device according to Note 8, wherein the video information includes one or more of face recognition information, object recognition information, gesture information, expression information, and posture information.

[0142] Note 10. The information processing device according to Note 8, wherein the audio information includes voiceprint recognition information and / or sound conversion information.

[0143] Supplement 11. The information processing device according to Supplement 8, wherein the three-dimensional information includes one or more of spatial position information, motion information, and distance information.

[0144] Note 12. The information processing device according to any one of Notes 2 to 11, wherein the processing circuit is configured to render the virtual space based on a trigger event generated by the real-time interaction, thereby obtaining an updated virtual space.

[0145] Note 13. The information processing device according to Note 12, wherein the triggering event is the generation of film and television special effects or online teaching special effects.

[0146] Note 14. The information processing device according to Note 12 or 13, wherein the processing circuit is further configured to display the updated virtual space instead of the virtual space on the display device that displays the virtual space, thereby obtaining an updated display image related to the triggering event.

[0147] Note 15. The information processing apparatus according to Note 14, wherein the display device comprises an LED screen.

[0148] Note 16. The information processing device according to Note 14 or 15, wherein the processing circuit is configured to enable the camera device to capture the predetermined area when the trigger event occurs and the image displayed by the display device, thereby obtaining a captured image.

[0149] Note 17. The information processing device according to Note 16, wherein the processing circuit is configured to determine the position and / or size of the rendering in the virtual space based on the position and / or posture relationship between the camera device and the display device.

[0150] Note 18. An information processing method comprising:

[0151] In response to a trigger based on at least one element in the real space, a trigger event is generated in the virtual space corresponding to the real space, thereby achieving real-time interaction between the real space and the virtual space.

[0152] Note 19. The information processing method according to Note 18, wherein a predetermined area in the real space including the at least one element is simulated to construct a simulation area, and the real-time interaction is realized based on the association between the simulation area and the virtual space.

[0153] Note 20. The information processing method according to Note 19, wherein the simulation processing is performed based on perception data obtained by at least one sensor sensing the predetermined area.

[0154] Note 21. The information processing method according to Note 20, wherein the at least one sensor includes one or more of a visible light sensor, an infrared sensor, an ultraviolet sensor, a distance perception sensor, an audio sensor, and a motion sensor.

[0155] Note 22. An information processing method according to any one of Notes 19 to 21, wherein at least one element included in the simulation area is analyzed, and the trigger is performed when the at least one element meets a predetermined trigger condition, and wherein the at least one element included in the simulation area corresponds to the at least one element included in the real space.

[0156] Note 23. The information processing method according to Note 22, wherein the at least one element includes one or more of an entity, background, light, and sound in the simulation area.

[0157] Note 24. The information processing method according to Note 22 or 23, wherein, based on element information related to the at least one element, it is determined whether the at least one element meets the predetermined trigger condition, thereby performing the triggering.

[0158] Note 25. The information processing method according to Note 24, wherein the element information includes one or more of video information, audio information and three-dimensional information related to the at least one element.

[0159] Note 26. The information processing method according to Note 25, wherein the video information includes one or more of face recognition information, object recognition information, gesture information, expression information, and posture information.

[0160] Note 27. The information processing method according to Note 25, wherein the audio information includes voiceprint recognition information and / or sound conversion information.

[0161] Note 28. The information processing method according to Note 25, wherein the three-dimensional information includes one or more of spatial position information, motion information, and distance information.

[0162] Note 29. The information processing method according to any one of Notes 19 to 28, wherein the virtual space is rendered based on a triggering event generated by the real-time interaction, thereby obtaining an updated virtual space.

[0163] Note 30. The information processing method according to Note 29, wherein the triggering event is the generation of film and television special effects or online teaching special effects.

[0164] Note 31. The information processing method according to Note 29 or 30, wherein the updated virtual space is displayed on a display device that displays the virtual space instead of the virtual space, thereby obtaining an updated display image related to the triggering event.

[0165] Note 32. The information processing method according to Note 31, wherein the display device includes an LED screen.

[0166] Note 33. The information processing method according to Note 31 or 32, wherein the camera device is configured to capture the predetermined area when the trigger event occurs and the image displayed by the display device, thereby obtaining a captured image.

[0167] Note 34. The information processing method according to Note 33, wherein the position and / or size of the rendering in the virtual space is determined based on the position and / or posture relationship between the camera device and the display device.

[0168] Note 35. A computer-readable storage medium having computer-executable instructions stored thereon, which, when executed, performs the information processing method according to any one of Notes 18 to 34.

Claims

1. An information processing device, comprising: The processing circuit is configured to: In response to a trigger based on at least one element in the real space, a trigger event is generated in the virtual space corresponding to the real space, thereby achieving real-time interaction between the real space and the virtual space.

2. The information processing device according to claim 1, wherein The processing circuit is configured to perform simulation processing on a predetermined area in the real space including the at least one element to construct a simulation area, and implement the real-time interaction based on an association between the simulation area and the virtual space.

3. The information processing device according to claim 2, wherein The processing circuit is configured to perform the simulation processing based on sensing data obtained by at least one sensor sensing the predetermined area.

4. The information processing device according to claim 3, wherein The at least one sensor includes one or more of a visible light sensor, an infrared sensor, an ultraviolet sensor, a distance sensing sensor, an audio sensor, and a motion sensor.

5. The information processing device according to any one of claims 2 to 4, wherein: The processing circuit is configured to analyze at least one element included in the simulation area and to perform the trigger when the at least one element meets a predetermined trigger condition, and wherein the at least one element included in the simulation area corresponds to the at least one element included in the real space. The information processing device according to claim 5 , wherein: The at least one element includes one or more of an entity, a background, a light, and a sound in the simulation area.

7. The information processing device according to claim 5 or 6, wherein: The processing circuit is configured to determine whether the at least one element satisfies the predetermined trigger condition based on element information related to the at least one element, thereby performing the triggering.

8. The information processing device according to claim 7, wherein The element information includes one or more of video information, audio information, and three-dimensional information related to the at least one element.

9. The information processing device according to claim 8, wherein The video information includes one or more of face recognition information, object recognition information, gesture information, expression information, and posture information.

10. The information processing device according to claim 8, wherein The audio information includes voiceprint recognition information and / or sound conversion information.

11. The information processing device according to claim 8, wherein The three-dimensional information includes one or more of spatial position information, motion information, and distance information.

12. The information processing apparatus according to any one of claims 2 to 11, wherein: The processing circuit is configured to render the virtual space based on the triggering event generated by the real-time interaction, thereby obtaining an updated virtual space.

13. The information processing device according to claim 12, wherein: The triggering event is to generate film and television special effects or online teaching special effects.

14. The information processing device according to claim 12 or 13, wherein: The processing circuit is further configured to display the updated virtual space on a display device that displays the virtual space instead of the virtual space, thereby obtaining an updated display image related to the trigger event.

15. The information processing device according to claim 14, wherein The display device includes an LED screen.

16. The information processing device according to claim 14 or 15, wherein: The processing circuit is configured to enable the camera device to capture the predetermined area when the trigger event occurs and the image displayed by the display device, thereby obtaining a captured image.

17. The information processing device according to claim 16, wherein: The processing circuit is configured to determine a position and / or size of the rendering in the virtual space based on a position and / or posture relationship between the camera device and the display device.

18. An information processing method, comprising: In response to a trigger based on at least one element in the real space, a trigger event is generated in the virtual space corresponding to the real space, thereby achieving real-time interaction between the real space and the virtual space.

19. The information processing method according to claim 18, wherein: A predetermined area in the real space including the at least one element is simulated to construct a simulation area, and the real-time interaction is achieved based on an association between the simulation area and the virtual space.

20. The information processing method according to claim 19, wherein: The simulation process is performed based on sensing data obtained by at least one sensor sensing the predetermined area.

21. The information processing method according to claim 20, wherein: The at least one sensor includes one or more of a visible light sensor, an infrared sensor, an ultraviolet sensor, a distance sensing sensor, an audio sensor, and a motion sensor.

22. The information processing method according to any one of claims 19 to 21, wherein: At least one element included in the simulation area is analyzed, and the trigger is performed when the at least one element meets a predetermined trigger condition, and wherein the at least one element included in the simulation area corresponds to the at least one element included in the real space.

23. The information processing method according to claim 22, wherein: The at least one element includes one or more of an entity, a background, a light, and a sound in the simulation area.

24. The information processing method according to claim 22 or 23, wherein: Based on element information related to the at least one element, it is determined whether the at least one element meets the predetermined trigger condition, thereby performing the triggering.

25. The information processing method according to claim 24, wherein: The element information includes one or more of video information, audio information, and three-dimensional information related to the at least one element.

26. The information processing method according to claim 25, wherein: The video information includes one or more of face recognition information, object recognition information, gesture information, expression information, and posture information.

27. The information processing method according to claim 25, wherein: The audio information includes voiceprint recognition information and / or sound conversion information.

28. The information processing method according to claim 25, wherein: The three-dimensional information includes one or more of spatial position information, motion information, and distance information.

29. The information processing method according to any one of claims 19 to 28, wherein: The virtual space is rendered based on the triggering event generated by the real-time interaction, thereby obtaining an updated virtual space.

30. The information processing method according to claim 29, wherein: The triggering event is to generate film and television special effects or online teaching special effects.

31. The information processing method according to claim 29 or 30, wherein: The updated virtual space is displayed on a display device that displays the virtual space instead of the virtual space, thereby obtaining an updated display image related to the triggering event.

32. The information processing method according to claim 31, wherein: The display device includes an LED screen.

33. The information processing method according to claim 31 or 32, wherein: The camera device is enabled to shoot the predetermined area when the trigger event occurs and the image displayed by the display device, thereby obtaining a shot image.

34. The information processing method according to claim 33, wherein: Based on the position and / or posture relationship between the camera device and the display device, the position and / or size of the rendering in the virtual space is determined. 35 . A computer-readable storage medium having computer-executable instructions stored thereon, wherein when the computer-executable instructions are executed, the information processing method according to claim 18 is executed.