Video recording method and device in virtual scene, storage medium and electronic device

By controlling a virtual camera to capture image data in a virtual scene and using an audio engine to process audio signals, the problem of sound and image mismatch in virtual reality video recording is solved, achieving a more realistic recording effect.

CN115604408BActive Publication Date: 2025-12-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211185076.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-12-05
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

In existing technologies, virtual reality video recording methods result in a mismatch between sound and visuals. The footage shot from a third-person perspective is inconsistent with the footage displayed on the headset, leading to a discrepancy between the sound and the actual sound.

Method used

By controlling a virtual camera to capture image data from the camera's perspective, and using an audio engine to spatially process the original audio signal based on the pose information of the virtual camera and the sound source, matching audio data is generated and synthesized to record video.

Benefits of technology

It achieves sound and visual matching in virtual scene video recording, improving the realism and immersion of the recorded video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115604408B_ABST
    Figure CN115604408B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video recording method and device in a virtual scene, a storage medium and an electronic device, and relates to the technical field of computers. The method comprises: controlling a first virtual camera to collect first picture data of the virtual scene under a camera view angle; and controlling an audio engine to perform spatialization processing on an original audio signal triggered by a sound-emitting body based on first pose information of the first virtual camera and second pose information of the sound-emitting body in the virtual scene, to obtain first audio data, and combining the first audio data and the first picture data to obtain a first recording video. In this way, the sound and the picture in the first recording video are matched, which conforms to the video recording in a real situation and greatly increases the authenticity of the video recorded in the virtual scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, in particular, to a video recording method and device in a virtual scene, a storage medium and an electronic device. BACKGROUND

[0002] With the rapid growth of the virtual reality (VR) market and users, the demand for VR recording by users is increasing, especially the users want to capture the scene of playing games. In the related art, the virtual scene is mainly captured by using a third-person perspective shooting method, a picture of the virtual scene is obtained, then the picture is synthesized with the sound output by the system play layer of the VR device to obtain a recorded video.

[0003] However, this video recording method can cause the sound in the video to be mismatched with the picture, that is, the picture captured by the third-person perspective is inconsistent with the picture displayed by the helmet, which causes the sound generated based on the picture displayed by the helmet to be inconsistent with the actual sound corresponding to the picture captured by the third-person perspective. SUMMARY

[0004] This summary is provided to introduce a selection of concepts, which will be described with greater detail in the detailed description below. This summary does not intend to identify key or essential features of the claimed technology, nor is it intended for use in determining the scope of the claimed technology.

[0005] In a first aspect, the present disclosure provides a video recording method in a virtual scene, comprising:

[0006] In response to a video recording instruction, determining a first virtual camera in the virtual scene;

[0007] In response to a target operation for the first virtual camera, controlling the first virtual camera to collect first picture data of the virtual scene in a camera perspective; and

[0008] Controlling an audio engine to process an original audio signal of a sound-emitting body based on first pose information of the first virtual camera and second pose information of the sound-emitting body in the virtual scene, to obtain first audio data;

[0009] Obtaining the first audio data generated by the audio engine, and obtaining the first picture data collected by the first virtual camera;

[0010] Based on the first picture data and the first audio data, obtaining a first recorded video.

[0011] In a second aspect, the present disclosure provides a video recording device in a virtual scene, comprising:

[0012] a creating module configured to determine a first virtual camera in the virtual scene in response to a video recording instruction;

[0013] a picture collecting module configured to control the first virtual camera to collect first picture data of the virtual scene under a camera view angle in response to a target operation on the first virtual camera;

[0014] an audio generating module configured to control an audio engine to process an original audio signal of a sound-emitting body based on first pose information of the first virtual camera and second pose information of the sound-emitting body in the virtual scene, to obtain first audio data;

[0015] an obtaining module configured to obtain the first audio data generated by the audio engine, and obtain the first picture data collected by the first virtual camera;

[0016] a video generating module configured to obtain a first recording video based on the first picture data and the first audio data.

[0017] In a third aspect, the present disclosure provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processing device to implement the steps of the method of the first aspect.

[0018] In a fourth aspect, the present disclosure provides an electronic device, comprising:

[0019] a storage device having a computer program stored thereon;

[0020] a processing device configured to execute the computer program in the storage device to implement the steps of the method of the first aspect.

[0021] According to the above technical solution, by controlling the first virtual camera to collect first picture data of the virtual scene under a camera view angle, and controlling the audio engine to perform spatialization processing on an original audio signal triggered by the sound-emitting body based on first pose information of the first virtual camera and second pose information of the sound-emitting body in the virtual scene, to obtain first audio data, and combining the first audio data and the first picture data to obtain a first recording video, the sound and the picture in the first recording video are matched, which conforms to the video recording in a real situation, and greatly increases the authenticity of the video recorded in the virtual scene.

[0022] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0023] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings. The same or similar components have the same or similar reference labels. It should be understood that the drawings are not necessarily to scale, with emphasis instead being placed upon illustrating the principles of the embodiments of the present disclosure. In the drawings:

[0024] Figure 1 is a flowchart of a method for video recording in a virtual scene according to some embodiments.

[0025] Figure 2 is a schematic diagram of an interface for controlling a first virtual camera according to some embodiments.

[0026] Figure 3 is a schematic diagram of an audio engine generating first audio data and second audio data according to some embodiments.

[0027] Figure 4 is a schematic diagram of recording a first recorded video according to some embodiments.

[0028] Figure 5 is a schematic diagram of recording a first recorded video according to yet some embodiments.

[0029] Figure 6 is a schematic diagram of a flow of audio data according to some embodiments.

[0030] Figure 7 is a flowchart of a method for video recording in a virtual scene according to some other embodiments.

[0031] Figure 8 is a schematic diagram of a module connection of a video recording apparatus in a virtual scene according to some embodiments.

[0032] Figure 9 is a schematic diagram of a structure of an electronic device according to some embodiments. DETAILED DESCRIPTION

[0033] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several embodiments of the present disclosure have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the present disclosure. It is to be understood that the drawings and descriptions are not to scale, with emphasis instead being placed upon illustrating the principles of the embodiments of the present disclosure. In the drawings:

[0034] It should be understood that each step recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.

[0035] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising, but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related definitions are given below.

[0036] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.

[0037] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise explicitly stated in the context, it should be understood as "one or more".

[0038] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0039] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.

[0040] For example, when recording a video, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.

[0041] As an optional but not limited implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.

[0042] It can be understood that the above notification and user authorization obtaining process is only illustrative, and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0043] At the same time, it can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0044] The embodiment of the present disclosure provides a video recording method and device in a virtual scene, a storage medium and an electronic device. The video recording method in the virtual scene can be used in the game field and the virtual reality technology field.

[0045] In the game field, the user can start the video recording function of the game application program on the electronic device, so as to record the posture of the game character operated by the user and the scene in the game scene during the process of playing the game by the user.

[0046] Among them, the game application program can be a three-dimensional game application program, which is a three-dimensional electronic game made on the basis of three-dimensional computer graphics. Compared with traditional two-dimensional games, it can bring players a more realistic game experience. Three-dimensional games refer to games made with three-dimensional technology, not to screens output in three dimensions.

[0047] In the virtual reality technology field, when the user wears a VR headset to play a VR application program, the user can start the video recording function in the VR application program on the electronic device, so as to record the posture of the virtual character taken by the user and the scene in the virtual scene during the process of playing the VR application program by the user.

[0048] Among them, the VR application program refers to an interactive application that provides immersive experience in a three-dimensional environment generated on a computer by comprehensively using a computer graphics system and various interfaces of reality and control, etc. The three-dimensional environment generated on a computer and interactable is called virtual scene (English full name: Virtual Environment, English abbreviation: VE). The VR application program can provide a human-computer interface for the user, so that the user can command the VR device installed with the VR application program and how the VR device provides information to the user.

[0049] Figure 1 is a flowchart of a video recording method in a virtual scene according to some embodiments. As shown in Figure 1As shown, the embodiment of the present disclosure provides a video recording method in a virtual scene, which can be executed by an electronic device, specifically, by a video recording device in a virtual scene, which can be implemented by software and / or hardware and configured in an electronic device. As shown in Figure 1 As shown, the method can include the following steps.

[0050] S110, in response to a video recording instruction, determining a first virtual camera in the virtual scene.

[0051] Here, the video recording instruction can be triggered by a user through a preset operation. For example, the user can trigger the video recording instruction by clicking a recording button in the user interface. For another example, the user can trigger the video recording instruction by clicking a hardware button of the electronic device, such as when the electronic device is a VR device, the user can trigger the video recording instruction by clicking a hardware button in the handle. Of course, the video recording instruction can also be triggered in other ways, such as through voice, action, etc., which will not be repeated here.

[0052] The electronic device determines the first virtual camera in response to the video recording instruction. The first virtual camera can be a virtual camera that has been created in the virtual scene or a virtual camera that has been recreated in the virtual scene. The first virtual camera is a camera assumed in the software of the electronic device and is a tool for representing a viewpoint in a three-dimensional virtual environment, which is used to record the picture data of the game scene in the virtual scene. That is, the first virtual camera is a camera assumed in the video recording program in the electronic device, through which the scene image can be recorded in the virtual scene. It should be understood that the first virtual camera can be represented by a model in the virtual scene, for example, the first virtual camera can be a camera model or a selfie stick model in the virtual scene.

[0053] In the embodiment of the present disclosure, the virtual scene can be a three-dimensional game scene or a virtual reality scene.

[0054] S120, in response to a target operation for the first virtual camera, controlling the first virtual camera to collect first picture data of the virtual scene at a camera view angle.

[0055] Here, the target operation for the first virtual camera can be an operation of moving or rotating the first virtual camera by the user. Through the target operation for the first virtual camera, the first virtual camera can be shot at different angles and distances in the virtual scene.

[0056] As some examples, the target operation for the first virtual camera can be triggered by the user through a virtual button in the user interface.Figure 2 is a schematic diagram of an interface for controlling a first virtual camera according to some embodiments. As shown in Figure 2 the user interface 200, the left side is a control area, and the right side is a virtual scene 201 in which a first virtual camera 202 is created. The control area of the user interface 200 includes an orientation control area 203 and a rotation control area 204. The orientation control area 203 is used to control the movement of the first virtual camera 202 in three directions, i.e., up and down, left and right, and forward and backward. The rotation control area 204 is used to control the rotation of the first virtual camera 202 in two directions, i.e., up and down, and left and right. It should be understood that the user can control the first virtual camera 202 to collect first picture data of the virtual scene 201 at different angles by touch operations on the orientation control area 203 and the rotation control area 204 of the user interface 200.

[0057] As another example, the target operation for the first virtual camera can be triggered by the user through the VR handle. When the user can start the video recording function in the VR application on the electronic device, the user can control the movement and rotation of the first virtual camera by operating the VR handle. For example, the pose of the VR handle is bound to the pose of the first virtual camera, and when the VR handle rotates, the first virtual camera synchronously rotates.

[0058] It should be noted that the camera view is the third-person shooting view of the first virtual camera in the virtual scene. The target operation of the user for the first virtual camera is used to adjust the shooting view of the first virtual camera, that is, to adjust the shooting angle of the camera view. The first picture data can refer to the picture material of the virtual scene presented in the camera view, for example, it can be a certain area in a three-dimensional game scene, or a certain area in a virtual reality scene.

[0059] S130, controlling an audio engine to process an original audio signal of the sound-emitting body based on the first pose information of the first virtual camera and the second pose information of the sound-emitting body in the virtual scene, to obtain first audio data.

[0060] Here, the first audio data is audio data obtained by the audio engine performing spatialization processing on an original audio signal emitted by a sound emitter in the virtual scene based on first pose information of a first virtual camera and second pose information of the sound emitter. The sound emitter refers to an object in the virtual scene that emits a specific sound effect. For example, if a pile of firewood in the virtual scene emits a crackling sound, there is a sound emitter at the location of the firewood, and the sound emitter emits the crackling sound. It should be understood that the sound emitter can be in an invisible state in the virtual scene, or the sound emitter only represents a sound event of an object playing a sound effect, and it does not necessarily mean a physical object. In addition, the sound emitter can also be referred to as an emitter in other terms.

[0061] In some embodiments, the audio engine can also process the original audio data based on the third pose information of the virtual character and the second pose information to generate second audio data while generating the first audio data.

[0062] The virtual character refers to a game character operated by a player, and the original audio signal refers to the most original audio file created by a sound designer using a tool similar to Audacity. In the audio engine, the virtual character operated by the player corresponds to a second listener, and the audio engine performs spatialization processing on the original audio signal triggered by the emitter based on the third pose information of the second listener and the second pose information of the emitter to obtain the second audio data. It should be noted that the pose information (including the first pose information, the second pose information, and the third pose information) involved in the embodiments of the present disclosure can refer to position information and attitude information, wherein the position information determines the distance between the second listener and the emitter, and the attitude information determines the orientation between the second listener and the emitter.

[0063] Exemplarily, the spatialization processing on the original audio signal includes distance processing and direction processing. The distance processing changes one or more of loudness, frequency, diffuseness, and focus of the original audio signal according to a distance parameter of the emitter relative to the second listener. The direction processing changes timbre and the like of the original audio signal according to an orientation parameter of the emitter relative to the second listener.

[0064] It should be understood that the second audio data after spatialization processing is actually a sound signal with six degrees of freedom of front and back, left and right, up and down, pitch, roll, and yaw.

[0065] The second audio data is used for output through the audio output device, and the second audio data is actually the sound heard by the player through the electronic device. For example, for a virtual reality scene, the second audio data is output through the audio output device of the VR headset, and is actually the sound heard by the wearer of the VR headset. For another example, when the audio output device of the electronic device used by the player is a headset, the second audio data is binauralized and output through the headset. For another example, when the audio output device configured by the electronic device is a separate loudspeaker, the second audio data is loudspeaker processed and output through the loudspeaker.

[0066] In the process of outputting the second audio data by the electronic device through the audio output device, the audio engine also simultaneously performs spatialization processing on the original audio signal based on the first pose information of the first virtual camera and the second pose information of the sound source in the virtual scene, to obtain the first audio data.

[0067] In the audio engine, a first listener is created, and the fourth pose information of the first listener is associated with the pose information of the first virtual camera, i.e., the fourth pose information of the first listener is equal to the first pose information of the first virtual camera. The audio engine performs spatialization processing on the original audio signal based on the fourth pose information and the second pose information of the first listener, to obtain the first audio data. It should be understood that the first audio data is actually equivalent to the sound heard by the first virtual camera at its current position and orientation.

[0068] Figure 3 FIG. 1 is a schematic diagram of an audio engine generating first audio data and second audio data according to some embodiments. As shown in FIG. 1, when recording a video in a virtual scene, the audio engine 300 has two audio streams. A first audio stream is that the audio engine 300 performs spatialization processing on the original audio signal based on the second pose information of the sound source corresponding to the original audio signal and the second pose information of the second listener, to obtain the second audio data. The second audio data is sent by the audio engine 300 to the audio playback system of the electronic device, and the second audio data is played back through the audio playback system to be output on the audio output device, to form the sound actually heard by the user in the virtual scene. Figure 3 A second audio stream is that the audio engine 300 performs spatialization processing on the original audio signal based on the second pose information of the sound source corresponding to the original audio signal and the fourth pose information of the first listener, to obtain the first audio data. The first audio data is not output through the audio playback system and the audio output device. That is, the first audio data is actually the sound heard by the first virtual camera at its current position and orientation, rather than the sound actually heard by the user in the virtual scene.

[0069] It is worth noting that, in this embodiment of the disclosure, the audio engine may be Wwise (a spatial engine).

[0070] S140, acquire the first audio data generated by the audio engine, and acquire the first image data captured by the first virtual camera.

[0071] Here, electronic devices can extract first audio data from the audio engine and first image data captured by the first virtual camera from the game engine.

[0072] S150, based on the first image data and the first audio data, obtain the first recorded video.

[0073] Here, the electronic device merges the first screen data and the first audio data according to the timeline of the first screen data and the first audio data to form the first recorded video.

[0074] It is worth noting that the sound in the first recorded video was captured from the camera's perspective of the first virtual camera, and the sound matches the video data in the first recorded video.

[0075] Figure 4 This is a schematic diagram illustrating the principle of recording a first recorded video, based on some embodiments. For example... Figure 4 As shown, in the virtual scene 401, the tanker truck 402 acts as a sound source, emitting a horn sound. The first virtual camera 403 is located to the left of the tanker truck 402, and the virtual character 404 is located to the right of the tanker truck 402. At this time, from the camera's perspective of the first virtual camera 403, the horn sound emitted by the tanker truck 402 should come from the right side of the first virtual camera 403. From the perspective of the virtual character 404, the horn sound emitted by the tanker truck 402 should come from the left side of the virtual character 404. Therefore, by using an audio engine based on the first pose information of the first virtual camera 403 and the second pose information of the sound source (tanker truck 402) in the virtual scene 401, the original audio signal (horn sound) triggered by the sound source is spatialized to obtain the first audio data, enabling the sound in the first recorded video to match the first frame data of the virtual scene 401 captured by the first virtual camera 403 from its camera's perspective. Moreover, users can still hear realistic sounds that match their field of vision. That is, the audio engine spatializes the original audio signal (horn sound) triggered by the sound source based on the third pose information of the virtual character 404 and the second pose information of the sound source (tanker truck 402) in the virtual scene 401, obtains the second audio data, and outputs the second audio data through the audio output device, so that users can hear realistic horn sounds that match their current pose.

[0076] Figure 5 Fig. 1 is a schematic diagram illustrating a principle of recording a first recorded video according to some embodiments. As shown, a user can create a selfie stick 102 (equivalent to a first virtual camera) in a virtual scene 101 while playing a virtual reality game, to record a video of a fire burning in the virtual scene 101 through the selfie stick 102. At this time, since the selfie stick 102 is closer to the fire than a virtual character operated by the user, the sound of the fire burning actually recorded by the selfie stick 102 should be louder than the sound of the fire burning heard by the user. The audio engine performs spatialization processing on the original audio signal (sound of the fire burning) triggered by the sound-emitting body (fire) based on the first pose information of the selfie stick 102 and the second pose information of the sound-emitting body (fire) in the virtual scene 101, to obtain first audio data, so that the sound in the first recorded video can match the first picture data of the virtual scene 101 collected by the selfie stick 102 in the camera view. Moreover, the user can still hear the real sound consistent with the visual angle range of the user, i.e., the audio engine performs spatialization processing on the original audio signal (sound of the fire burning) triggered by the sound-emitting body based on the third pose information of the virtual character (not shown in Fig. 1) and the second pose information of the sound-emitting body (fire) in the virtual scene 101, to obtain second audio data, and outputs the second audio data through an audio output device, so that the user can hear the real sound of the fire burning consistent with the current pose of the user. Figure 5 Figure 5

[0077] Thus, by controlling the first virtual camera to collect first picture data of the virtual scene in the camera view, and controlling the audio engine to perform spatialization processing on the original audio signal triggered by the sound-emitting body based on the first pose information of the first virtual camera and the second pose information of the sound-emitting body in the virtual scene, to obtain first audio data, and combining the first audio data and the first picture data to obtain the first recorded video, the sound in the first recorded video is matched with the picture, which is consistent with the video recording in the real situation, greatly increasing the authenticity of the video recorded in the virtual scene.

[0078] In some implementable embodiments, in S140, obtaining the first audio data generated by the audio engine can include: sampling the first audio data through a first audio output management plug-in, to obtain an audio sampling signal in pulse code modulation format.

[0079] ​​Here, in the audio engine, a first audio output management plug-in (Audio Route) can be configured to sample the first audio data generated by the audio engine to obtain an audio sample signal in a pulse code modulation (PCM) format, so that the electronic device can generate the first recorded video based on the sampled audio sample signal and the obtained first picture data.

[0080] It is worth noting that the first audio output management plug-in samples the first audio data by actually outputting the first audio data generated by the audio engine to the game engine of the electronic device in the form of an audio stream, so that the first audio data is merged with the first picture data in the game engine to generate the first recorded video.

[0081] In some implementable embodiments, in S150, the first recorded video is obtained based on the first picture data and the first audio data, including:

[0082] The audio sample signal output by the first audio output management plug-in is received by the game engine, and the audio sample signal and the first picture data are merged to obtain the first recorded video.

[0083] Here, the first audio output management plug-in can send the sampled audio sample signal to the game engine, and the game engine can merge the audio sample signal and the first picture data according to the time axis of the audio sample signal and the first picture data to obtain the first recorded video. The first picture data can be obtained by the electronic device through the game engine, i.e., the game engine creates a first virtual camera to collect the first picture data of the virtual scene.

[0084] In the game engine, a video recording module can be created, which is configured to control the start and end of video recording, receive the audio sample signal, and merge the audio sample signal with the first picture data to output the first recorded video.

[0085] In some implementable embodiments, the electronic device can also control the first audio output management plug-in to mute the audio sample signal so that the audio sample signal is not sent to the audio playback system.

[0086] Here, the mute processing is not to adjust the loudness of the audio sample signal sampled by the first audio output management plug-in to 0, but to directly adjust the signal amount of the audio sample signal to 0, so that the first audio output management plug-in does not send the audio sample signal to the audio playback system.

[0087] It should be understood that when the first audio output management plug-in does not send the audio sample signal to the audio playback system, the audio output device of the electronic device will not output the audio sample signal. At this time, in the user's hearing, the user will only hear the second audio data and will not hear the first audio data.

[0088] Thus, by muting the audio sample signal, the audio output device of the electronic device can be prevented from simultaneously outputting the first audio data and the second audio data, and sound interference will not occur. In this case, the sound in the first recorded video recorded by the first virtual camera is the first audio data, and the sound actually heard by the user is the second audio data.

[0089] In some other implementable embodiments, the electronic device can also control the audio output device to output the first audio data and control the audio output device not to output the second audio data in response to an audio playback instruction, wherein the second audio data is obtained by processing the original audio signal based on the third pose information of the virtual character and the second pose information by the audio engine.

[0090] Here, the audio playback instruction is triggered by the user, and the user can select the audio signal that he or she wants to play in the user interface. For example, selecting to play the first audio data, selecting to play the second audio data, or selecting to play the first audio data and the second audio data simultaneously. Outputting the first audio data by the audio output device means that the first audio data is played through the audio playback system of the electronic device and the audio output device, forming the sound heard by the user.

[0091] Among them, the first audio output management plug-in can send the sampled audio sample signal to the audio playback system so that the audio playback system plays back the audio sample signal and plays it through the audio output device. In the case of outputting the first audio data, the audio output device can be controlled not to output the second audio data so that the user can only hear the first audio data and will not hear the second audio data.

[0092] It is worth noting that the second audio data has been described in detail in the above embodiments and will not be repeated here.

[0093] Thus, by outputting the first audio data, the user can experience the real sound effect of the recorded video when recording the video, so as to adjust the recording angle according to the sound effect and obtain a better first recorded video. In addition, the first audio output management plug-in sends the PCM format audio sample signal to the game engine, which is a frame-by-frame audio stream data, rather than an entire audio file. Based on this, the game engine can synchronize the frame audio sample signal with the corresponding first picture data to generate a video every time it receives a frame of audio sample signal, thereby improving the generation speed of the first recorded video.

[0094] The following will be described in detail with reference to the accompanying drawings Figure 6 The above embodiments will be described in detail.

[0095] Figure 6 is a schematic diagram of the flow direction of audio data according to some embodiments. As Figure 6 indicated, the audio engine 600 performs spatialization processing on the original audio signal triggered by the sound source based on the first pose information of the first listener and the second pose information of the sound source in the virtual scene to obtain the first audio data, and is configured to process the original audio signal based on the third pose information of the second listener and the second pose information of the sound source to obtain the second audio data.

[0096] Among them, the audio engine 300 outputs the second audio data to the audio playback system, and the audio playback system performs playback on the second audio data, and outputs the second audio data through the audio output device to form the sound heard by the user.

[0097] For the first audio data, the audio engine 300 samples the first audio data through the first audio output management plug-in to form the PCM format audio sample signal. The first audio output management plug-in performs muting processing on the audio sample signal, so that the first audio output management plug-in does not send the audio sample signal to the audio playback system. It should be understood that if the first audio output management plug-in does not perform muting processing on the audio sample signal, the audio sample signal is necessarily sent to the audio playback system. That is, the sound signal generated by the audio engine 300 will be necessarily output to the audio playback system for playing. In addition, the first audio output management plug-in also sends the audio sample signal to the game engine, and the game engine combines the received audio sample signal and the first picture data collected by the first virtual camera to generate the first recorded video.

[0098] Figure 7 is a flowchart of a video recording method in a virtual scene according to some embodiments. As Figure 7 indicated, the video recording method in the virtual scene can include the following steps:

[0099] S710, the audio engine splits the third audio data into at least one audio stream according to the sound effect type of the audio included in the third audio data, wherein each audio stream includes audio of one sound effect type, and the third audio data is obtained by processing the original audio signal of the sound-emitting body in the virtual scene.

[0100] Here, the third audio data can be the first audio data or the second audio data of the above-mentioned embodiments.

[0101] The sound effect type of the audio can include background music, skill sound effect, speaking sound, loudspeaker sound, vehicle sound, etc. The audio engine splits the third audio data into at least one audio stream based on the sound effect type of the audio included in the third audio data. For example, the third audio data includes background music and speaking sound, and the audio engine can split the third audio data into two audio streams of background music and speaking sound, each of which includes audio corresponding to one sound effect type.

[0102] S720, from the at least one audio stream, a target audio stream belonging to a target sound effect type is obtained.

[0103] Here, the target sound effect type refers to the sound effect type that the user wants to record. For example, when the user needs to record the dialogue in the game, a target audio stream belonging to the sound effect type of speaking sound can be obtained from the at least one audio stream. The target sound effect type can be one or multiple, which is set according to the user's needs.

[0104] In some embodiments, the second audio output management plug-in can be used to sample an audio stream belonging to the target sound effect type from the at least one audio stream to obtain a target audio stream in pulse code modulation format.

[0105] In the audio engine, a second audio output management plug-in can be provided, which is used to sample an audio stream belonging to the target sound effect type from the at least one audio stream to obtain a target audio stream in pulse code modulation format.

[0106] It should be understood that the functions and principles of the second audio output management plug-in have been described in detail in the above-mentioned embodiments with respect to the first audio output management plug-in, and will not be repeated here.

[0107] S730, based on the target audio stream and second picture data, a second recording video is generated, wherein the second picture data is the picture data of the virtual scene collected by the second virtual camera.

[0108] Here, the electronic device can synchronize the target audio stream and the second picture data according to a timeline of the target audio stream and the second picture data, and generate the second recording video through the game engine. Since the target audio stream is an audio stream corresponding to the target sound effect type, the second recording video actually includes the sound that the user wants to record and does not include the sound that the user does not want to record.

[0109] For example, when the user needs to record a video of a skill release, and the user wants the sound effect of the skill release to be more obvious, the user can set a target audio stream that samples the sound effect of the skill. In this way, the generated second recording video includes the sound effect of the skill and does not include other background sound, environmental sound, or the like.

[0110] For another example, when the user finds that the background music in the virtual scene is copyrighted music, the user can choose not to record the background music, i.e., the selected target sound effect type does not include the background music. In this way, the generated second recording video does not include the background music of the copyrighted music, thereby avoiding copyright risks when sharing or uploading the video.

[0111] In some embodiments, the game engine can generate the second recording video based on the target audio stream in the pulse code modulation format and the second picture data. The second picture data can be collected by a second virtual camera, and the acquisition principle is consistent with that of the first picture data, which will not be described herein again.

[0112] The second audio output management plug-in sends the target audio stream in the PCM format to the game engine, which is a frame-by-frame audio stream data, rather than an entire audio file. Based on this, the game engine can synchronize the target audio stream and the corresponding second picture data to generate a video every time a frame of the target audio stream is received, thereby improving the generation speed of the second recording video.

[0113] In the above embodiments, the electronic device can determine a second virtual camera in the virtual scene in response to a recording instruction, control the second virtual camera to collect second picture data of the virtual scene, control an audio engine to generate third audio data based on fourth pose information of a virtual character and fifth pose information of a sound source, and control the audio engine to split the third audio data into at least one audio stream according to a sound effect type of audio included in the third audio data, obtain a target audio stream belonging to a target sound effect type from the at least one audio stream, and generate a second recording video based on the target audio stream and the second picture data.

[0114] In this way, by splitting the third audio data and obtaining the target audio stream belonging to the target sound effect type from the split multiple audio streams, the user can customize the sound in the virtual scene that needs to be recorded, and copyright risks can also be avoided when the user shares the video.

[0115] Figure 8 is a module connection diagram of a video recording device in a virtual scene according to some embodiments. As shown in Figure 8 the disclosure embodiment provides a video recording device in a virtual scene, the device 800 comprises:

[0116] The creation module 801 is configured to determine a first virtual camera in the virtual scene in response to a video recording instruction.

[0117] The picture acquisition module 802 is configured to control the first virtual camera to acquire first picture data of the virtual scene under the camera view angle in response to a target operation for the first virtual camera.

[0118] The audio generation module 803 is configured to control an audio engine to process an original audio signal of a sound-emitting body based on first pose information of the first virtual camera and second pose information of the sound-emitting body in the virtual scene, and obtain first audio data.

[0119] The acquisition module 804 is configured to acquire the first audio data generated by the audio engine, and acquire the first picture data acquired by the first virtual camera.

[0120] The video generation module 805 is configured to obtain a first recording video based on the first picture data and the first audio data.

[0121] Optionally, the acquisition module 804 is specifically configured to:

[0122] sample the first audio data through a first audio output management plug-in to obtain an audio sample signal in pulse code modulation format.

[0123] Optionally, the video generation module 805 is specifically configured to:

[0124] receive the audio sample signal output by the first audio output management plug-in through a game engine, and combine the audio sample signal and the first picture data to obtain the first recording video.

[0125] Optionally, the device 800 further comprises:

[0126] The mute module is configured to control the first audio output management plug-in to perform mute processing on the audio sample signal, so that the audio sample signal is not sent to an audio playback system.

[0127] Optionally, the device 800 further comprises:

[0128] The output module is configured to, in response to the audio playing instruction, control the audio output device to output the first audio data and control the audio output device not to output second audio data, wherein the second audio data is obtained by processing the original audio signal based on third pose information of the virtual role and the second pose information by the audio engine.

[0129] Optionally, the apparatus 800 further includes:

[0130] The audio splitting module is configured to control the audio engine to split the third audio data into at least one audio stream according to the sound effect type of the audio included in the third audio data, wherein each of the at least one audio stream includes audio of one sound effect type, and the third audio data is obtained by processing the original audio signal of the sound-emitting body in the virtual scene.

[0131] The extraction module is configured to obtain a target audio stream belonging to a target sound effect type from the at least one audio stream.

[0132] The recording module is configured to generate a second recorded video based on the target audio stream and second picture data, wherein the second picture data is picture data of the virtual scene collected by a second virtual camera.

[0133] Optionally, the extraction module is specifically configured to:

[0134] sample an audio stream belonging to the target sound effect type from the at least one audio stream through a second audio output management plug-in to obtain a target audio stream in pulse code modulation format.

[0135] The recording module is specifically configured to:

[0136] generate a second recorded video based on the target audio stream in pulse code modulation format and the second picture data.

[0137] As to the apparatus 800 in the above embodiments, the specific manners in which various modules perform operations have been described in details in the embodiments of the method, and thus will not be described in details here.

[0138] Reference is made below to Figure 9 which shows a structural schematic diagram of an electronic device 900 suitable for use to implement embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle terminal (e.g., a vehicle navigation terminal), a VR device, and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. Figure 9The electronic device shown is merely an example and should not bring any limitation to the function and scope of use of the embodiments of the present disclosure.

[0139] As shown in Figure 9 The electronic device 900 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 902 or programs loaded into a random access memory (RAM) 903 from a storage device 908. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0140] In general, the following devices can be connected to the I / O interface 905: input devices 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 908 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 909. The communication devices 909 can allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 9 The electronic device 900 is shown with various devices, but it is understood that all of the devices shown are not required to be implemented or present. More or fewer devices can alternatively be implemented or present.

[0141] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 909, or installed from the storage devices 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0142] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable storage medium or carried by a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take various forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can be used to carry or store a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.

[0143] In some embodiments, the electronic device can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications of any form or medium (e.g., a communications network). Examples of communications networks include local area networks ("LAN"), wide area networks ("WAN"), internetworks (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future developed networks.

[0144] The aforementioned computer-readable medium can be contained in the aforementioned electronic device; or can exist separately without being assembled into the electronic device.

[0145] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: in response to a video recording instruction, determine a first virtual camera in a virtual scene; in response to a target operation for the first virtual camera, control the first virtual camera to collect first picture data of the virtual scene under a camera view angle; and control an audio engine to process an original audio signal of a sounder in the virtual scene based on first pose information of the first virtual camera and second pose information of the sounder in the virtual scene, to obtain first audio data; acquire the first audio data generated by the audio engine, and acquire the first picture data collected by the first virtual camera; and based on the first picture data and the first audio data, obtain a first recorded video.

[0146] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages, including an object-oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0147] The flow and block diagrams in the drawings show architectural, functional, and operational architectures of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0148] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0149] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0150] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of a program of a processor, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0151] The above description is merely exemplary of the present disclosure and the application of the principles thereof and the scope of the disclosure is not limited to the specific embodiments described herein, but only by the claims that follow. It will be readily apparent to those skilled in the art that varying substitutions and modifications can be made to the embodiments described invented without departing from the scope of the present disclosure. Accordingly, the disclosure is not to be restricted as illustrated and described herein, but is amenable to alterations and modifications, and the application is intended to include all such alterations and modifications as fall within the scope of the application. Thus, it is intended that the application cover the modifications and variations of this application provided they come within the scope of the appended claims and their equivalents.

[0152] Moreover, while operations can be depicted in a particular, serial order, this should not be understood as requiring or implying that such operations be performed in the order illustrated, or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, while certain features can be described in only one or several embodiments, it should be understood that they can be combined in any manner in other embodiments as well. Similarly, while operations can be described as being performed by a single device, it should be understood that such operations can be performed by multiple devices acting in concert.

[0153] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the various modules perform the operations has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.

Claims

1. A video recording method in a virtual scene, characterized by, The method comprises: in response to a video recording instruction, determining a first virtual camera in a virtual scene; in response to a target operation on the first virtual camera, controlling the first virtual camera to collect first picture data of the virtual scene under a camera view angle; and controlling an audio engine to process original audio signals of a sound-emitting body in the virtual scene based on first pose information of the first virtual camera and second pose information of the sound-emitting body in the virtual scene, to obtain first audio data; the audio engine processes the original audio signals based on third pose information of a virtual role and the second pose information to generate second audio data for output through an audio output device while generating the first audio data; obtaining the first audio data generated by the audio engine and the first picture data collected by the first virtual camera; based on the first picture data and the first audio data, obtaining a first recorded video.

2. The method of claim 1, wherein, The method further comprises: sampling the first audio data through a first audio output management plug-in to obtain an audio sampling signal in pulse code modulation format.

3. The method of claim 2, wherein, The method further comprises: receiving, through a game engine, the audio sampling signal output by the first audio output management plug-in, and merging the audio sampling signal and the first picture data to obtain the first recorded video.

4. The method of claim 2, wherein, The method further comprises: controlling the first audio output management plug-in to mute the audio sampling signal so that the audio sampling signal is not sent to an audio playback system.

5. The method of claim 1, wherein, The method further comprises: in response to an audio playback instruction, controlling the audio output device to output the first audio data and controlling the audio output device not to output second audio data, wherein the second audio data is obtained by processing the original audio signals based on third pose information of a virtual role and the second pose information by the audio engine.

6. The method of claim 1, wherein, The method further comprises: controlling the audio engine to split third audio data into at least one audio stream according to the sound effect type of the audio included in the third audio data, wherein each audio stream includes audio of one sound effect type, and the third audio data is obtained by processing original audio signals of a sound-emitting body in the virtual scene; from the at least one audio stream, obtaining a target audio stream belonging to a target sound effect type; based on the target audio stream and second picture data, generating a second recorded video, wherein the second picture data is picture data of the virtual scene collected by a second virtual camera.

7. The method of claim 6, wherein, The method further comprises: sampling, through a second audio output management plug-in, an audio stream belonging to the target sound effect type from the at least one audio stream to obtain a target audio stream in pulse code modulation format; the method further comprises: based on the target audio stream and second picture data, generating a second recorded video, wherein the second picture data is picture data of the virtual scene collected by a second virtual camera. generate a second recorded video based on the target audio stream in the pulse code modulation format and the second picture data.

8. A video recording apparatus in a virtual scene, characterized by comprising: The method comprises the steps of: creating a module configured to determine a first virtual camera in a virtual scene in response to a video recording instruction; a picture acquisition module configured to control the first virtual camera to acquire first picture data of the virtual scene under a camera view angle in response to a target operation for the first virtual camera; an audio generation module configured to control an audio engine to process an original audio signal of a sound-emitting body based on first pose information of the first virtual camera and second pose information of the sound-emitting body in the virtual scene, to obtain first audio data; the audio engine processes the original audio signal based on third pose information of a virtual role and the second pose information at the same time of generating the first audio data, to generate second audio data for output through an audio output device; an acquisition module configured to acquire the first audio data generated by the audio engine, and acquire the first picture data acquired by the first virtual camera; a video generation module configured to obtain a first recorded video based on the first picture data and the first audio data.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by a processing device to implement the steps of the method of any one of claims 1 to 7.

10. An electronic device, comprising: The method comprises the steps of: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multiple listener cloud render with enhanced instant replay

    US20180332422A1