Video recording method, device and electronic equipment
Through the wireless communication connection between the first device and the second device, image and sound information are obtained and integrated to generate multimedia files, which solves the problem that external devices such as mobile phones cannot synchronously process the virtual content of AR devices and realizes the realism of augmented reality.
Patent Information
- Application Number
- CN202211709899.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-12-29
AI Technical Summary
Existing external devices such as mobile phones cannot synchronize the virtual content of AR devices when recording audio or video.
Through wireless communication connection between the first device and the second device, information to be fused is obtained, including the first image information and the first sound information displayed by the second device, the second image information and the second sound information displayed by the first device, and fusion processing is performed to generate a fused multimedia file.
The fusion and superposition of images and sounds from different devices is achieved. The obtained multimedia file not only has real recorded images and sounds, but also has superimposed virtual images and sounds. The first device can simultaneously synchronize the virtual content of the second device during recording and videotaping, thereby enhancing the realism of augmented reality.
Smart Images

Figure CN115988154B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and specifically relates to a video recording method, device and electronic equipment. Background Art
[0002] Currently, with the emergence of more and more consumer electronic products, consumers maintain an open and inclusive attitude towards new forms and technologies of smart terminal products. New smart terminal devices such as augmented reality (AR) are developing very rapidly, and consumers are also quickly adapting to such products, and even regard them as necessities of life.
[0003] In AR technology, a micro-projection system projects virtual information, such as text and images, onto an optical element. This information is then transmitted to the human eye through reflection and total internal reflection. The real-world scene can then be directly viewed through the optical element, allowing the user to see an "overlap" of virtual and real life, thus achieving augmented reality. In existing technologies, because the virtual information and real-world scenes displayed by AR devices exist independently, recording or videotaping the AR device's display content requires the AR device itself to process the virtual content. External devices, such as mobile phones, cannot synchronize the recording and videotaping of the AR device's virtual content. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a recording method, device and electronic device to solve the problem that existing external devices such as mobile phones cannot synchronize the virtual content of AR devices during recording and video recording.
[0005] In order to solve the above technical problems, this application is implemented as follows:
[0006] In a first aspect, an embodiment of the present application provides a recording method, which is applied to a first device, wherein the first device is wirelessly connected to a second device, and the recording method includes:
[0007] Acquire information to be fused, where the information to be fused includes: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information;
[0008] Performing fusion processing on the information to be fused to obtain a fused first multimedia file;
[0009] An image corresponding to the first multimedia file is displayed.
[0010] In a second aspect, an embodiment of the present application provides a recording method, which is applied to a second device, the second device being wirelessly connected to the first device, the recording method comprising:
[0011] Acquire information to be fused, where the information to be fused includes: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information;
[0012] The information to be fused is subjected to fusion processing to obtain a fused second multimedia file.
[0013] In a third aspect, an embodiment of the present application provides a video recording device, which is applied to a first device, wherein the first device is wirelessly connected to a second device, and the video recording device includes:
[0014] A first acquisition module is configured to acquire information to be fused, the information to be fused comprising: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information;
[0015] A first fusion module is used to perform fusion processing on the information to be fused to obtain a fused first multimedia file;
[0016] The display module is configured to display an image corresponding to the first multimedia file.
[0017] In a fourth aspect, an embodiment of the present application provides a video recording device, which is applied to a second device, the second device being wirelessly connected to the first device, and the video recording device comprising:
[0018] a second acquisition module, configured to acquire information to be fused, the information to be fused comprising: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information;
[0019] The second fusion module is used to perform fusion processing on the information to be fused to obtain a fused second multimedia file.
[0020] In a fifth aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect or the second aspect are implemented.
[0021] In a sixth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.
[0022] In the seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the method described in the first aspect or the second aspect.
[0023] In an embodiment of the present application, the first device fuses the first image information and the corresponding first sound information displayed by the second device with the second image information and the corresponding sound information displayed by the first device to obtain a first multimedia file, and displays the image in the first multimedia file, thereby realizing the fusion and superposition of images and sounds of different devices. The obtained multimedia file not only has real recorded images and sounds, but also has superimposed virtual images and sounds. The first device can simultaneously synchronize the virtual content of the second device when recording audio and video, thereby strengthening the interaction between the second device and the first device, and making augmented reality more realistic. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is one of the flowcharts of the video recording method according to an embodiment of the present application;
[0025] Figure 2 is a schematic diagram of the working principle of the second device in an embodiment of the present application;
[0026] Figure 3 This is one of the structural diagrams of the first device and the second device in the embodiment of the present application;
[0027] Figure 4 This is the second flow chart of the video recording method according to an embodiment of the present application;
[0028] Figure 5 This is the second structural diagram of the first device and the second device in the embodiment of the present application;
[0029] Figure 6 This is the third flow chart of the video recording method according to an embodiment of the present application;
[0030] Figure 7 This is the third structural diagram of the first device and the second device in the embodiment of the present application;
[0031] Figure 8 This is the fourth flow chart of the video recording method according to an embodiment of the present application;
[0032] Figure 9 This is the fourth structural diagram of the first device and the second device in the embodiment of the present application;
[0033] Figure 10 This is the fifth flow chart of the video recording method according to an embodiment of the present application;
[0034] Figure 11This is the sixth flow chart of the video recording method according to an embodiment of the present application;
[0035] Figure 12 This is one of the structural diagrams of the video recording device according to an embodiment of the present application;
[0036] Figure 13 This is the second structural diagram of the video recording device according to an embodiment of the present application;
[0037] Figure 14 This is one of the structural diagrams of the electronic device according to the embodiment of the present application;
[0038] Figure 15 This is the second structural diagram of the electronic device according to the embodiment of the present application. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0040] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0041] The image processing method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0042] like Figure 1 As shown, an embodiment of the present application provides a recording method, which is applied to a first device, wherein the first device is wirelessly connected to a second device, and the recording method includes:
[0043] Step 101: Acquire information to be fused, where the information to be fused includes: first image information displayed by a second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information.
[0044] The first device displays the second image information, and the second image information has corresponding second sound information; the second device displays the first image information, and the first image information has corresponding first sound information. The first device and the second device can exchange information to obtain the above image information and sound information.
[0045] The first sound information and the second sound information may be sound signals, and the first image information and the second image information may be image signals.
[0046] Optionally, the first device may be a mobile terminal (eg, a mobile phone terminal), the second device may be an AR device (eg, AR glasses), and the first device and the second device are wirelessly connected.
[0047] In the embodiment of the present application, the image information displayed by the second device and the first device may be played image information or recorded image information. For example, the second device may display the first image information by playing a virtual image and its corresponding first sound information, and the first device may display the second image information by photographing an image of the real environment, that is, displaying the second image information of the real environment on the screen of the first device. The first image information and the corresponding first sound information, and the second image information and the corresponding second sound information are the information that needs to be fused.
[0048] Take the second device as an example, AR glasses, such as Figure 2 As shown, the micro-projection system projects virtual information such as text and images onto optical elements, and then transmits the virtual information to the human eye through reflection and total reflection; for the images in the real scene, they can directly enter the human eye through the optical elements, and the user can see the "overlap" of virtual information and the real scene, thereby realizing augmented reality.
[0049] Step 102: performing fusion processing on the information to be fused to obtain a fused first multimedia file;
[0050] Step 103: Display the image corresponding to the first multimedia file.
[0051] In this embodiment, after obtaining the information to be fused, the first device may superimpose and fuse the first image information, the second image information, and the corresponding first and second sound information into a multimedia file. Specifically, the first multimedia file includes not only the fused image information but also the sound information corresponding to the image. The first device may display the image of the fused first multimedia file and, optionally, may also play the sound of the first multimedia file.
[0052] In an embodiment of the present application, the first device fuses the first image information and the corresponding first sound information displayed by the second device with the second image information and the corresponding sound information displayed by the first device to obtain a first multimedia file, and displays the image in the first multimedia file, thereby realizing the fusion and superposition of images and sounds of different devices. The obtained multimedia file not only has real recorded images and sounds, but also has superimposed virtual images and sounds. The first device can simultaneously synchronize the virtual content of the second device when recording audio and video, thereby strengthening the interaction between the second device and the first device, and making augmented reality more realistic.
[0053] The method further includes sending the first multimedia file to the second device.
[0054] After obtaining the merged first multimedia file, the first device may send the file to the second device, and the second device may play and / or store the first multimedia file.
[0055] As an optional embodiment, obtaining the information to be fused includes:
[0056] receiving first image information sent by the second device;
[0057] Acquire the first sound information and the second sound information by one of the following methods:
[0058] Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by the microphone of the first device;
[0059] Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the second device are received;
[0060] Method 3: When the second device plays the first sound information in headphone mode, the second sound information is collected by the microphone of the first device, and the first sound information sent by the second device is received;
[0061] Method 4: When the second device plays the first sound information in headphone mode, the first sound information and the second sound information sent by the second device are received.
[0062] In this embodiment, since the second image information is displayed by the first device, it does not need to be acquired again. The first image information is displayed by the second device, so the second device needs to send the first image information to the first device.
[0063] When the first device obtains the first sound information and the second sound information, they may be recorded by the microphone of the first device, or may be recorded by the microphone of the second device and sent to the first device.
[0064] Among them, when the first device or the second device records sound, if the first sound information of the second device is played in external speaker mode, the first sound information and the second sound information can be recorded through the microphone of the first device; or, the first sound information and the second sound information can be recorded by the second device, and the second device can send the recorded first sound information and the second sound information to the first device.
[0065] If one of the sound information is played through the headphone mode, only the other sound information can be recorded. For example: the second device plays the virtual first image information, and its corresponding first sound information is played through the headphone mode. The first device records and shoots the real environment and the corresponding second sound information, and displays the second image information of the real environment on the screen of the first device. Then, the first sound information needs to be sent from the second device to the first device. Alternatively, the first sound information of the second device is played through the headphone mode, and the second sound information can be recorded by the second device, and the first sound information stored by itself and the recorded second sound information are sent to the first device.
[0066] As an optional embodiment, the method further includes: receiving relative position information sent by the second device, where the relative position information is relative position information between the first image information and the second image information;
[0067] The fusing the information to be fused includes: fusing the information to be fused according to the relative position information.
[0068] In this embodiment, the relative display position may be a relative coordinate. The first device may fuse the first image information and the second image information into one according to the relative coordinates of the two, thereby realizing image fusion between different devices.
[0069] The second device has an image position recognition function, which means that it can recognize the relative position between the first image information and the second image information. For example, the second device determines the relative display position of the first image information and the second image information in the following manner: the second device has a camera that can capture the picture in the real environment, analyzes and processes the real environment information transmitted by the shooting module, and combines intelligent recognition technology with simultaneous localization and mapping (SLAM) technology to recognize that the user wants to present the virtual picture in the relative display position of the real environment. The second device can then know the relative display position of the virtual image information (i.e., the first image information) in the real picture (i.e., the second image information), and the second device can superimpose the virtual image on the real picture according to the relative display position; or, the second device can also send the virtual image and the relative display position to the first device, and the first device superimposes the virtual image on the real picture. The relative display position in this embodiment can be a relative coordinate information.
[0070] If the first device does not have an image position recognition function, the second device determines the phase display position of the first image information and the second image information, and sends the relative display position to the first device. For example, if a mobile terminal does not have an image position recognition function, that is, it cannot determine where the user wants the AR device's virtual image to be displayed on the real image, the AR device can determine the relative display position and send it to the mobile terminal. The mobile terminal can then overlay the virtual image on the real image based on the relative display position.
[0071] As an optional embodiment, the first sound information has a corresponding first time code, the second sound information has a corresponding second time code, the first image information has a corresponding third time code, and the second image information has a corresponding fourth time code;
[0072] The fusion processing of the information to be fused includes: matching the first time code with the third time code in time position, and matching the second time code with the fourth time code in time position; fusing the first sound information with the first image information, fusing the second sound information with the second image information, and fusing the first image information with the second image information according to the matched time position, to obtain the first multimedia file.
[0073] In this embodiment, the first and second devices can add time codes to the displayed image information and recorded audio information. The time codes can be clock codes that run in frames. This time code can track the recording (or display) time of the video and audio. This time code information is then stored as metadata in the digital file. When performing information fusion, the time codes on the image and the time codes on the audio can be matched to achieve audio and video synchronization, resulting in an image file that combines images from different devices and the corresponding audio.
[0074] The following takes the second device being AR glasses and the first device being a mobile phone terminal as an example to illustrate the implementation process of the recording method of the embodiment of the present application.
[0075] 1. Take the example of an AR glasses playing the first sound information corresponding to the virtual image in an external speaker mode, and the AR glasses having an image position recognition function.
[0076] like Figure 3 As shown, the mobile phone terminal 310 is used to record the second image information 320 and sound information 321 of the real environment. The sound information 321 may include the second sound information of the real scene and the first sound information corresponding to the virtual image played in the external speaker mode of the AR glasses; the AR glasses 340 play the virtual image information 330, and play the first sound information 342 in the virtual image information 330 in the external speaker mode through the speaker 341 of the AR glasses.
[0077] The AR glasses 340 are provided with headphones 343, and the headphones 343 are provided with a microphone, which can be used to record the sound of the real environment and the sound information in the virtual image information played in the external speaker mode; the AR glasses 340 and the mobile phone terminal 310 are connected through a predetermined communication method 350, for example, through Bluetooth. After the communication between the AR glasses 340 and the mobile phone terminal 310 is established, information can be transmitted to each other through the predetermined communication method 350.
[0078] This embodiment may include two implementation methods, in which the mobile phone terminal 310 records sound information and the AR glasses 340 records sound information, which are described below with examples.
[0079] (1) The microphone of the mobile phone terminal obtains real sound information, and the AR glasses play virtual sound information in the external speaker mode. The sound signal is large and can be picked up by the microphone of the mobile phone terminal. The AR device transmits the acquired image information and sound information to the mobile phone terminal in real time. The data interaction between the AR device and the mobile phone terminal is realized through Bluetooth technology. The AR device has a camera that can capture the image in the real environment, analyze and process the real environment information transmitted by the shooting module, and combine intelligent recognition technology with SLAM technology to identify the relative position where the user wants to present the virtual image in the real environment. The AR device can send this relative position to the mobile phone terminal. Based on this relative position, the mobile phone terminal superimposes the virtual image sent by the AR device on the real scene image, and superimposes the sound information recorded by the mobile phone terminal on the image. The mobile phone terminal realizes the superposition and fusion of the virtual image, the real scene image, and the recorded sound information, and saves the fused image file, achieving perfect synchronization between the image information, sound information and the saved file during recording. Alternatively, the image fusion process is implemented by an AR device, which superimposes virtual image information on the display image sent by the mobile phone terminal according to the relative position, and fuses the sound information and image information recorded by the mobile phone terminal to complete the fusion processing and transmission of real and virtual sound information.
[0080] (2) The headset microphone of the AR glasses acquires real sound information, and the AR device transmits the acquired image information to the mobile phone terminal. The data interaction between the AR device and the mobile phone terminal is realized through Bluetooth technology. The AR device can analyze the relative position relationship and sound position relationship between the virtual image information and the real scene image, and maintain the original time correspondence relationship between the real scene image information and the sound information, and the original time correspondence relationship between the virtual image information and the virtual sound information. The AR device can send the relative position of the virtual image and the real scene image to the mobile phone terminal, and send the recorded sound information to the mobile phone terminal. The mobile phone terminal can superimpose the virtual image on the real scene image based on the relative position relationship between the virtual image information and the real scene image, and superimpose the sound information recorded by the AR glasses on the image. Alternatively, the image fusion process is realized by the AR device, and the mobile phone terminal sends the second image information to the AR device. The AR device superimposes and fuses the virtual image, the real scene image, and the recorded sound information based on the relative position information, and saves the fused image file.
[0081] Taking the mobile phone terminal to realize information fusion as an example, the implementation process of the recording method is as follows: Figure 4 Shown, including:
[0082] Step 0: A Bluetooth connection is established between the mobile terminal and the AR glasses. The user starts recording the display environment through the mobile terminal. The AR glasses play the virtual image and the corresponding sound information of the virtual image in the external speaker mode.
[0083] Step 1: Select the microphone of the mobile terminal or the headset microphone of the AR glasses to record the sound information.
[0084] Step 21: If sound information is recorded through the microphone of the mobile terminal, when the AR glasses play the virtual image and the sound information of the virtual image, the mobile terminal records the real environment sound and the sound information corresponding to the virtual image played by the AR glasses.
[0085] Step 31: The AR glasses transmit virtual image information to the mobile terminal in real time.
[0086] Step 41: The mobile phone terminal receives the image information transmitted by the AR glasses.
[0087] Step 51: The mobile phone terminal fuses the virtual image and the real image according to their relative positions to ensure synchronization with the information seen by the human eye.
[0088] Step 61: The mobile terminal matches the fused image information with the sound information in time to ensure the synchronization of the picture and sound. It should be noted that steps 51 and 61 can be performed in a sequential order or in an unsequential order (i.e., performing sound and image fusion simultaneously). This example only uses the example of performing image fusion first and then sound fusion to illustrate.
[0089] Step 71: The mobile phone terminal sends the processed multimedia file to the AR glasses, which can preview and save the image file.
[0090] Alternatively, step 22: if sound information is recorded through the headset microphone of the AR glasses, then when the AR glasses play the virtual image and the sound information of the virtual image, the headset microphone of the AR glasses records the real environment sound and the sound information corresponding to the virtual image played by the AR glasses;
[0091] Step 32: The AR glasses transmit the recorded real image information, the sound information recorded by the microphone, and the virtual image information in real time.
[0092] Step 42: The mobile phone terminal receives the image information and sound information transmitted by the AR glasses.
[0093] Step 52: The mobile phone terminal fuses the virtual image and the real image according to their relative positions to ensure synchronization with the information seen by the human eye.
[0094] Step 62: The mobile phone terminal matches the fused image information with the sound information in time to ensure the synchronization of the picture and the sound.
[0095] Step 72: The mobile phone terminal sends the processed multimedia file to the AR glasses, which can preview and save the image file.
[0096] 2. Take the example of AR glasses playing the first sound information corresponding to the virtual image through the headphone mode, and the AR glasses having the image position recognition function.
[0097] like Figure 5 As shown, the mobile phone terminal 510 is used to record the second image information 520 and the second sound information 521 of the real environment; the AR glasses 540 play the virtual image information 530, and play the first sound information 542 in the virtual image information through the earphones 541, wherein the sound leakage is very small in the earphone mode, and the microphone of the user's mobile phone cannot record the first sound information in the virtual information; at the same time, the earphones 541 are provided with a microphone, which can be used to record the real environment sound, and the earphones 541 can be set to the transparent mode to play the real environment sound and the virtual sound information together.
[0098] The AR glasses 540 and the mobile phone terminal 510 are connected via a predetermined communication method 550, such as Bluetooth. After the communication between the AR glasses 540 and the mobile phone terminal 510 is established, information can be transmitted to each other via the predetermined communication method 550.
[0099] This embodiment may include two implementation methods, in which the mobile phone terminal 510 records sound information and the AR glasses 540 records sound information, which are described below with examples.
[0100] (A) The mobile phone terminal microphone captures the sounds in the real environment, and the AR glasses transmit the virtual image and the corresponding sound information to the mobile phone terminal in real time. The mobile phone terminal can superimpose the virtual image on the real scene image based on the relative position of the virtual image and the real scene image (which the AR device recognizes and sends to the mobile phone terminal), and superimpose the sound information recorded by the mobile phone terminal on the image. Alternatively, the mobile phone terminal sends the real image information and the sound information corresponding to the real image to the AR device, and the AR glasses realize the superposition and fusion of the virtual image, the real scene image, and the recorded sound information, and save the fused image file, completing the recording and fusion of the sound information in the AR virtual information and the sound information in the real environment.
[0101] (B) The headset microphone on the AR glasses captures the sound information of the real environment and sends it to the mobile terminal. The mobile terminal superimposes the virtual image on the real scene image based on the relative position of the virtual image information and the real scene image, and matches the sound information recorded by the AR glasses with the image information in time to ensure the synchronization of sound and picture. Alternatively, the mobile terminal sends the real image information to the AR glasses, which realize the superposition and fusion of the virtual image, the real scene image, and the recorded sound information, completing the fusion processing and transmission of the real and virtual sound information during the video recording.
[0102] Taking the first device to realize information fusion, the first device is a mobile phone terminal and the second device is AR glasses as an example, the implementation process of the recording method is as follows: Figure 6 including:
[0103] Step a: A Bluetooth connection is established between the mobile terminal and the AR glasses. The user starts recording the display environment through the mobile terminal. The AR glasses play the virtual image and the sound information corresponding to the virtual image through the headphone mode. The headphones can be set to transparent mode.
[0104] Step b: Select the microphone of the mobile terminal to record the sound information or the headset microphone of the AR glasses to record the sound information.
[0105] Step c1: If the sound information is recorded through the microphone of the mobile phone terminal, when the AR glasses play the virtual image and the sound information of the virtual image, the mobile phone terminal records the real environment sound.
[0106] Step d1: The AR glasses transmit virtual images and corresponding sound information to the mobile terminal in real time.
[0107] Step e1: The mobile phone terminal receives the virtual image information and sound information transmitted by the AR glasses, as well as the relative position of the virtual image and the real image.
[0108] Step f1: The mobile terminal fuses the virtual image and the real image based on their relative positions to ensure synchronization with the information seen by the human eye.
[0109] Step g1: The mobile phone terminal matches the fused image information with the sound information (the sound of the real environment recorded by the mobile phone terminal and the sound information corresponding to the virtual image on the AR glasses end) in time to ensure the synchronization of the picture and the sound.
[0110] Step h1: The mobile terminal sends the processed multimedia file to the AR glasses, which can preview and save the image file.
[0111] Alternatively, step c2: if the sound information is recorded by the earphone microphone of the AR glasses, then when the AR glasses play the virtual image and the sound information of the virtual image, the earphone microphone of the AR glasses records the sound information of the real environment;
[0112] Step d2: The AR glasses transmit virtual image information and corresponding sound information, as well as recorded sound information, to the mobile phone terminal in real time.
[0113] Step e2: The mobile phone terminal receives the image information and sound information transmitted by the AR glasses.
[0114] Step f2: The mobile terminal fuses the virtual image and the real image according to their relative positions to ensure synchronization with the information seen by the human eye.
[0115] Step g2: The mobile phone terminal matches the fused image information with the sound information in time to ensure the synchronization of the picture and the sound.
[0116] It should be noted that step f2 and step g2 may have a certain order or may not have a certain order (ie, the fusion of sound and image is performed simultaneously). Here, the example of performing image fusion first and then performing sound fusion is used for explanation.
[0117] Step h2: The mobile terminal sends the processed multimedia file to the AR glasses, which can preview and save the image file.
[0118] In this embodiment, the virtual image and virtual sound information of the AR device are combined with the real image and real sound information recorded by the mobile terminal, which can enhance the interaction between the AR device and the mobile terminal, making the augmented reality more realistic and meeting the user's needs for taking pictures in any background.
[0119] As an optional embodiment, the method further includes:
[0120] receiving a first input from a user;
[0121] In response to the first input, determining that the object corresponding to the first input is an object in the first image information, or determining that the object corresponding to the first input is an object in the second image information;
[0122] If the object corresponding to the first input is an object in the first image information, amplifying the first sound information or reducing the second sound information;
[0123] When the object corresponding to the first input is an object in the second image information, the second sound information is amplified, or the first sound information is reduced.
[0124] In this embodiment, the first input is used to achieve directional recording. A user can make the first input for image information in which they want to emphasize the sound. The first device can then amplify the sound corresponding to the corresponding image information to achieve directional recording. The first input can be a predetermined gesture operation, such as a finger tap or double-tap on the image information display location, and the like, without limitation.
[0125] For example: the first device detects the user's touch screen operation on the screen, and can determine whether the operation is for a virtual image or a real scene image based on the touch screen position. If the user performs a touch screen operation on the relative display position corresponding to the virtual image, the sound information corresponding to the virtual image is amplified (or the sound information of the real image is reduced or suppressed). If the user performs a touch screen operation on a position other than the relative display position corresponding to the virtual image, the sound information corresponding to the real scene image is amplified (or the sound information corresponding to the virtual image is reduced or suppressed).
[0126] The first image information is virtual image information transmitted by the second device. If the user's first input to the object in the first image information is detected, the first sound information corresponding to the first image information is amplified or the second sound information is reduced; the second image information is real image information captured by the first device. If the user's first input to the second image information is detected, the sound corresponding to the second image information needs to be amplified or the first sound information needs to be reduced.
[0127] In this embodiment, if sound information is recorded through the first device and the first sound information is amplified, in order to highlight the first sound information, the first device can suppress the second sound information during recording; or, if sound information is recorded through the first device and the second device is instructed to amplify the first sound information, in order to highlight the first sound information, the first device can suppress the second sound information during recording.
[0128] As an optional embodiment, when the first sound information and / or the second sound information is recorded through a microphone of the second device, the method further includes:
[0129] Sending instruction information to the second device;
[0130] Among them, if the first sound information is amplified, the indication information is used to instruct the second device to suppress the second sound information when recording the first sound information and the second sound information; if the second sound information is amplified, the indication information is used to instruct the second device to suppress the first sound information when recording the first sound information and the second sound information.
[0131] In this embodiment, if sound information is recorded through the second device and the first sound information is amplified, in order to highlight the first sound information, the first device can send an instruction message to the second device, instructing the second device to suppress the second sound information during recording; or, if sound information is recorded through the second device and the second sound information is amplified, in order to highlight the second sound information, the first device can send a second instruction message to the second device, instructing the second device to suppress the first sound information during recording.
[0132] It should be noted that in the embodiments of the present application, the amplification and sound suppression of the first or second sound information can be performed during the sound recording process of the first or second device. The recorded sound information is the sound information after the sound directional amplification and suppression process. When the first device fuses the image information and the sound information, the sound information after the sound directional amplification and suppression process is also used. The image fusion and sound fusion processes are not described in detail here.
[0133] Taking the example where the first device is a mobile phone terminal and the second device is AR glasses, the following example illustrates the implementation process of directional recording.
[0134] 1. Take the example of AR glasses playing the first sound information corresponding to the virtual image through the headphone mode, and the AR glasses having the image position recognition function.
[0135] like Figure 7 As shown, a mobile phone terminal 710 is used to record second image information 720 and second sound information 721 of a real environment; AR glasses 740 play virtual image information 730 and play first sound information 742 in the virtual image information through headphones 741. The AR glasses 740 and mobile phone terminal 710 are connected via a predetermined communication method 750, such as Bluetooth. The AR glasses 740 have a built-in camera 770.
[0136] The user can operate the virtual image information by focusing on the virtual image information through an operation gesture 760, or can operate the real image by focusing on the real image information through an operation gesture 761. The mobile terminal 710 can detect the user's input on the virtual image or the real image.
[0137] In this embodiment, when the user needs to focus on a certain picture, he or she can click on the corresponding picture information with a finger. When the mobile phone terminal 710 detects the user's action, it is considered that the user is focusing. According to the position of the finger on the picture during focusing, it can be determined whether the user is focusing on the virtual picture information or the real picture information. When the mobile phone terminal 710 detects that the user is focusing on the virtual picture 760, the mobile phone terminal 710 amplifies the virtual sound signal corresponding to the virtual image information, and at the same time suppresses the real environment sound signal transmitted by the mobile phone terminal (instruction information can be sent to the AR glasses), and then fuses the sound and image information. The specific fusion process is not described here.
[0138] When the mobile phone terminal 710 detects that the user has performed a focus operation 761 on a real image, the size and phase relationship of the sound signal is recorded through the microphone of the mobile phone terminal, the sound signal recorded by the mobile phone terminal is directionally amplified, and the sound signal in the remaining area is suppressed, thereby completing the directional function of the sound signal in the real environment, and then the sound and image information are fused. The specific fusion process will not be repeated here.
[0139] Optionally, the user's input can also be detected by the second device. For example, the AR glasses have a camera that can monitor the user's gesture operations, and can track the user's hand movements through intelligent recognition technology or SLAM technology to determine whether the user's first input is for virtual picture information or real picture information, thereby amplifying or suppressing the corresponding sound information. The specific implementation process is similar to the directional recording process of the first device and will not be repeated here.
[0140] This embodiment may include two implementation methods: performing the first input on a virtual image and performing the first input on a real image. Figure 8 As shown, taking the mobile terminal detecting the user's first input as an example, the process includes:
[0141] Step 1: Establish a Bluetooth connection between the mobile terminal and the AR glasses, and the mobile terminal or AR glasses start recording.
[0142] Step 2: The mobile terminal detects whether there is a focus operation (ie, the first input).
[0143] Step 31: Detect a focus operation on the virtual image.
[0144] Step 41: The mobile terminal amplifies the sound signal corresponding to the virtual image and suppresses the real environment sound signal transmitted by the AR device.
[0145] or
[0146] Step 32: Detect a focus operation on the real image.
[0147] Step 42: The mobile terminal directionally amplifies the sound signal in the area recorded by the mobile phone camera, while suppressing the sound signal in other areas. Optionally, the recorded sound information is transmitted to the AR glasses.
[0148] Step 43: The AR glasses may also amplify the real-world sound signal transmitted by the mobile phone terminal, while suppressing the sound signal corresponding to the virtual image.
[0149] Step 5: Perform fusion processing on the image information and the sound information. The fusion processing process will not be described in detail here.
[0150] 2. Take the example of an AR glasses playing the first sound information corresponding to the virtual image in an external speaker mode, and the AR glasses having an image position recognition function.
[0151] like Figure 9 As shown, a mobile phone terminal 910 is used to record real image information 920 and second sound information 921 of a real environment; AR glasses 940 play virtual image information 930 and play first sound information 942 in the virtual image information in an external speaker mode through the AR glasses' speakers 941. The AR glasses 940 and the mobile phone terminal 910 are connected via a predetermined communication method 950, such as Bluetooth.
[0152] The user can operate the virtual image information by focusing on the virtual image information through an operation gesture 960, or can operate the real image by focusing on the real image information through an operation gesture 961. The mobile terminal 910 detects the user's operation gesture, or the AR glasses are provided with a camera 970 for detecting the user's operation gesture.
[0153] The AR glasses 940 are provided with earphones 943, which are provided with a microphone, which can be used to record the sound of the real environment and the sound information in the virtual image information played in the external speaker mode; the AR glasses 940 and the mobile phone terminal 910 are connected via a predetermined communication method 950.
[0154] When the mobile phone terminal 910 or the camera 970 on the AR glasses detects that the user has performed a focus operation 960 on the virtual image, the size and phase of the virtual sound signal picked up by the AR glasses headset microphone and the mobile phone terminal microphone are different. Through processing by a directional recording algorithm, the direction of the sound source of the virtual sound signal can be located, and the sound signal in that direction can be amplified, while suppressing the sound signals in other directions, and then the sound and image information are fused. The specific fusion process is not repeated here.
[0155] When the mobile phone terminal 910 or the camera 970 on the AR glasses detects that the user has performed a focus operation 961 on the real image, the size and phase relationship of the sound signal recorded by the mobile phone terminal microphone is used to directionally amplify the sound signal in the recording area of the mobile phone camera, while suppressing the sound signal in the remaining area, completing the directional function of the sound signal in the real environment. The mobile phone terminal can transmit the processed sound signal to the AR glasses. The AR glasses amplify the real environment sound signal transmitted by the mobile phone terminal, while suppressing the virtual sound signal corresponding to the virtual image information, and then fuse the sound and image information. The specific fusion process is not repeated here.
[0156] This embodiment may include two implementation methods: performing the first input on a virtual image and performing the first input on a real image. Figure 10 As shown, taking the mobile terminal detecting the user's first input as an example, the process includes:
[0157] Step a: Establish a Bluetooth connection between the mobile phone terminal and the AR glasses, and the mobile phone terminal or the AR glasses starts recording.
[0158] Step b: The mobile phone terminal detects whether there is a focus operation (ie, the first input).
[0159] Step c1: detecting a focus operation on a virtual image.
[0160] Step d1: The mobile phone terminal locates the direction of the source of the virtual sound signal, amplifies the sound signal in that direction, and suppresses the sound signals in other directions.
[0161] or
[0162] Step c2: detecting a focus operation on the real image.
[0163] Step d2: The mobile phone terminal directionally amplifies the sound signal in the area recorded by the mobile phone camera, while suppressing the sound signal in other areas, and can transmit the recorded sound information to the AR glasses.
[0164] Step d3: The AR glasses amplify the real-world sound signal transmitted by the mobile phone terminal, while suppressing the virtual sound signal corresponding to the virtual image;
[0165] Step e: perform fusion processing on the image information and the sound information. The fusion processing process will not be described in detail here.
[0166] In this embodiment, the user makes a first input to the virtual image or the real image to achieve directional reception of the sound information corresponding to the virtual image in the AR device or the sound information of the real environment, thereby achieving directional reception of the sound of the video file, which can improve the user's video recording experience.
[0167] In an embodiment of the present application, the first device fuses the first image information and the corresponding first sound information displayed by the second device with the second image information and the corresponding sound information displayed by the first device to obtain a first multimedia file, and displays the image in the first multimedia file, thereby realizing the fusion and superposition of images and sounds of different devices. The obtained multimedia file not only has real recorded images and sounds, but also has superimposed virtual images and sounds. The first device can simultaneously synchronize the virtual content of the second device when recording audio and video, thereby strengthening the interaction between the second device and the first device, and making augmented reality more realistic.
[0168] like Figure 11 As shown, an embodiment of the present application further provides a recording method, which is applied to a second device, the second device being wirelessly connected to the first device, and the recording method includes:
[0169] Step 110: Acquire information to be fused, where the information to be fused includes: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information;
[0170] Step 111: perform fusion processing on the information to be fused to obtain a fused second multimedia file.
[0171] In this embodiment, the second device displays the first image information, and the first image information has corresponding first sound information; the first device displays the second image information, and the second image information has corresponding second sound information. The second device and the first device can exchange information.
[0172] The first device transmits the displayed second image information to the second device. After the second device obtains the second image information and the corresponding second sound information, the first image information and the corresponding first sound information, it can fuse the first image information and the second image information, the first sound information and the second sound information together to obtain a complete second multimedia file. After obtaining the fuser, the second device can send the multimedia file to the first device, and the first device can play and / or store the image file.
[0173] In an embodiment of the present application, a second device fuses the first image information and corresponding first sound information displayed by the second device with the second image information and corresponding sound information displayed by the first device to obtain a second multimedia file, thereby achieving the fusion and superposition of images and sounds from different devices to obtain a fused multimedia file. The fused multimedia file not only has real recorded images and sounds, but also has superimposed virtual images and sounds, thereby enhancing the interaction between the second device and the first device and making the augmented reality more realistic.
[0174] Optionally, the second device is an AR device (such as AR glasses), the first device is a mobile terminal (such as a mobile phone terminal), and the second device is wirelessly connected to the first device.
[0175] Optionally, the obtaining of information to be fused includes:
[0176] receiving second image information sent by the first device;
[0177] Acquire the first sound information and the second sound information, or acquire the second sound information, in one of the following ways:
[0178] Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the microphone of the first device are received;
[0179] Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by a microphone of the second device;
[0180] Method three: When the second device plays the first sound information in headphone mode, receiving the second sound information collected and sent by the first device;
[0181] Method 4: When the second device plays the first sound information in headphone mode, the second sound information is collected through the microphone of the second device.
[0182] When the second device obtains the first sound information and the second sound information, the first sound information and the second sound information may be recorded by the microphone of the second device, or may be recorded by the microphone of the first device and sent to the second device.
[0183] Among them, when the first device or the second device records sound, if the first sound information of the second device is played in external speaker mode, the first sound information and the second sound information can be recorded by the microphone of the first device and sent to the second device; or, the first sound information and the second sound information can also be recorded by the second device.
[0184] If one of the sound information is played through the headphone mode, only the other sound information can be recorded. For example: the second device plays the virtual first image information, and its corresponding first sound information is played through the headphone mode. The first device records and shoots the real environment and the corresponding second sound information and sends it to the second device. The first sound information can be obtained by the second device from its own storage. Alternatively, the first sound information of the second device is played through the headphone mode, and the second device can record the second sound information and obtain the first sound information stored in its own storage.
[0185] Optionally, the method further includes: determining relative position information of the first image information and the second image information; and the fusing the information to be fused includes: fusing the information to be fused according to the relative position information.
[0186] In this embodiment, the relative display position may be a relative coordinate. The second device may fuse the first image information and the second image information into one according to the relative coordinates of the first image information and the second image information, thereby realizing image fusion between different devices.
[0187] The second device has an image position recognition function, and the image position recognition function refers to the ability to recognize the relative position between the first image information and the second image information. For example: the way in which the second device determines the relative display position of the first image information and the second image information includes: the second device has a camera that can capture the picture in the real environment, analyzes and processes the real environment information transmitted by the shooting module, and combines with SLAM technology to recognize that the user wants to present the virtual picture in the relative display position of the real environment. Then, the second device can know the relative display position of the virtual image information (i.e., the first image information) in the real picture (i.e., the second image information), and the second device can superimpose the virtual image on the real picture according to the relative display position.
[0188] As an optional embodiment, the first sound information has a corresponding first time code, the second sound information has a corresponding second time code, the first image information has a corresponding third time code, and the second image information has a corresponding fourth time code;
[0189] The fusion processing of the information to be fused includes: matching the first time code with the third time code in time position, and matching the second time code with the fourth time code in time position; fusing the first sound information with the first image information, fusing the second sound information with the second image information, and fusing the first image information with the second image information according to the matched time position, to obtain the second multimedia file.
[0190] In this embodiment, the first and second devices can add time codes to the displayed image information and recorded audio information. The time codes can be clock codes that run in frames. This time code can track the recording (or display) time of the video and audio. This time code information is then stored as metadata in the digital file. When performing information fusion, the time codes on the image and the time codes on the audio can be matched to achieve audio and video synchronization, resulting in an image file that combines images from different devices and the corresponding audio.
[0191] The following example illustrates the process of image and sound fusion achieved by the second device (taking AR glasses as an example).
[0192] (1) The microphone of the mobile phone terminal obtains real sound information, and the AR glasses play virtual sound information in the external speaker mode. The sound signal is large and can be picked up by the microphone of the mobile phone terminal. The mobile phone terminal transmits the acquired image information and sound information to the AR glasses in real time. The data interaction between the AR glasses and the mobile phone terminal is realized through Bluetooth technology. The AR glasses have a camera that can capture the images in the real environment, analyze and process the real environment information transmitted by the shooting module, and combine intelligent recognition technology with SLAM technology to identify the relative position where the user wants to present the virtual image in the real environment. Based on the relative position, the AR glasses can superimpose the virtual image on the real image sent by the mobile phone terminal and fuse it with the sound information to achieve the superposition and fusion of the virtual image, the real scene image, and the recorded sound information, and save the fused image file to achieve perfect synchronization between the image information, sound information and the saved file during recording.
[0193] (2) The headset microphone of the AR glasses acquires real sound information, and the mobile phone terminal transmits the acquired image information to the AR glasses. The data interaction between the AR glasses and the mobile phone terminal is realized through Bluetooth technology. The AR glasses can analyze the relative position relationship and sound position relationship between the virtual image information and the real scene image, and maintain the original time correspondence between the real scene image information and sound information, and the original time correspondence between the virtual image information and the virtual sound information. The mobile phone terminal sends the second image information to the AR glasses, and the AR glasses superimpose and fuse the virtual image, the real scene image, and the recorded sound information according to the relative position information, and save the fused image file.
[0194] The process of the second device implementing the fusion of the first image information, the second image information, the first sound information and the second sound information is similar to the fusion process of the first device. The only difference is that the first device sends the image information to the second device, and the second device performs the fusion process. For the specific fusion process, please refer to the method embodiment of the first device and will not be repeated here.
[0195] Optionally, the method further includes: receiving a second input from a user;
[0196] In response to the second input, determining that the object corresponding to the second input is an object in the first image information, or determining that the object corresponding to the second input is an object in the second image information;
[0197] If the object corresponding to the second input is an object in the first image information, amplifying the first sound information, or reducing the second sound information;
[0198] When the object corresponding to the second input is an object in the second image information, the second sound information is amplified, or the first sound information is reduced.
[0199] In this embodiment, the second input is used to achieve directional recording. The user can make the second input for image information in which they want to emphasize the sound. The second device can then amplify the sound corresponding to the corresponding image information, achieving directional recording. The second input can be a predetermined gesture, such as sliding a finger in a preset pattern, etc., which is not limited here.
[0200] The second device detects the user's gesture operation and can determine whether the operation is directed at a virtual image or a real scene image based on the gesture operation. If the gesture operation is directed at a virtual image, the sound information corresponding to the virtual image is amplified (or the sound information of the real image is reduced or suppressed); if the gesture operation is directed at a real image, the sound information corresponding to the real scene image is amplified (or the sound information corresponding to the virtual image is reduced or suppressed).
[0201] Optionally, when the first sound information and the second sound information are recorded through a microphone of the first device, the method further includes:
[0202] Sending instruction information to the first device;
[0203] Among them, if the first sound information is amplified, the indication information is used to instruct the mobile phone terminal to suppress the second sound information when recording the first sound information and the second sound information; if the second sound information is amplified, the indication information is used to instruct the second device to suppress the first sound information when recording the first sound information and the second sound information.
[0204] In this embodiment, if sound information is recorded through the first device and the second sound information is amplified, in order to highlight the second sound information, the second device can send an instruction message to the first device, instructing the first device to suppress the first sound information during recording; or, if sound information is recorded through the first device and the second device amplifies the first sound information, in order to highlight the first sound information, the second device can send an instruction message to the first device, instructing the first device to suppress the second sound information during recording.
[0205] It should be noted that in the embodiments of the present application, the amplification and sound suppression of the first or second sound information can be performed during the sound recording process of the first or second device. The recorded sound information is the sound information after the sound directional amplification and suppression process. When the second device fuses the image information and the sound information, the sound information after the sound directional amplification and suppression process is also used. The image fusion and sound fusion processes are not described in detail here.
[0206] In an embodiment of the present application, a second device fuses the first image information and corresponding first sound information displayed by the second device with the second image information and corresponding sound information displayed by the first device to obtain a second multimedia file, thereby achieving the fusion and superposition of images and sounds from different devices to obtain a fused multimedia file. The fused multimedia file not only has real recorded images and sounds, but also has superimposed virtual images and sounds, which enhances the interaction between the second device and the mobile phone terminal and makes the augmented reality more realistic.
[0207] It should be noted that the recording method provided in the embodiment of the present application can be executed by a recording device or a control module in the recording device for the recording method. In the embodiment of the present application, the recording device provided in the embodiment of the present application is described by taking the recording method executed by the recording device as an example.
[0208] like Figure 12 As shown, an embodiment of the present application further provides a video recording device 1200, which is applied to a first device, wherein the first device is wirelessly connected to a second device, and the video recording device 1200 includes:
[0209] A first acquisition module 1210 is configured to acquire information to be fused, the information to be fused including: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information;
[0210] A first fusion module 1220 is configured to perform fusion processing on the information to be fused to obtain a fused first multimedia file;
[0211] The display module 1230 is configured to display an image corresponding to the first multimedia file.
[0212] Optionally, the first acquisition module is specifically configured to:
[0213] receiving first image information sent by the second device;
[0214] Acquire the first sound information and the second sound information by one of the following methods:
[0215] Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by the microphone of the first device;
[0216] Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the second device are received;
[0217] Method 3: When the second device plays the first sound information in headphone mode, the second sound information is collected by the microphone of the first device, and the first sound information sent by the second device is received;
[0218] Method 4: When the second device plays the first sound information in headphone mode, the first sound information and the second sound information sent by the second device are received.
[0219] Optionally, the device further includes:
[0220] a first receiving module, configured to receive relative position information sent by the second device, where the relative position information is relative position information between the first image information and the second image information;
[0221] The first fusion module is specifically configured to perform fusion processing on the information to be fused according to the relative position information.
[0222] Optionally, the device further includes:
[0223] A second receiving module, configured to receive a first input from a user;
[0224] a first response module, configured to, in response to the first input, determine that the object corresponding to the first input is an object in the first image information, or determine that the object corresponding to the first input is an object in the second image information;
[0225] If the object corresponding to the first input is an object in the first image information, amplifying the first sound information or reducing the second sound information;
[0226] When the object corresponding to the first input is an object in the second image information, the second sound information is amplified, or the first sound information is reduced.
[0227] In an embodiment of the present application, the first device fuses the first image information and the corresponding first sound information displayed by the second device with the second image information and the corresponding sound information displayed by the first device to obtain a first multimedia file, and displays the image in the first multimedia file, thereby realizing the fusion and superposition of images and sounds of different devices. The obtained multimedia file not only has real recorded images and sounds, but also has superimposed virtual images and sounds. The first device can simultaneously synchronize the virtual content of the second device when recording audio and video, thereby strengthening the interaction between the second device and the mobile phone terminal, making augmented reality more realistic.
[0228] like Figure 13As shown, the embodiment of the present application further provides a video recording device 1300, which is applied to a second device, and the second device is wirelessly connected to the first device. The video recording device 1300 includes:
[0229] The second acquisition module 1310 is configured to acquire information to be fused, where the information to be fused includes: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information;
[0230] The second fusion module 1320 is configured to perform fusion processing on the information to be fused to obtain a fused second multimedia file.
[0231] Optionally, the second acquisition module is specifically configured to:
[0232] receiving second image information sent by the first device;
[0233] Acquire the first sound information and the second sound information, or acquire the second sound information, in one of the following ways:
[0234] Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the microphone of the first device are received;
[0235] Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by a microphone of the second device;
[0236] Method three: When the second device plays the first sound information in headphone mode, receiving the second sound information collected and sent by the first device;
[0237] Method 4: When the second device plays the first sound information in headphone mode, the second sound information is collected through the microphone of the second device.
[0238] Optionally, the device further includes:
[0239] a first determining module, configured to determine relative position information of the first image information and the second image information;
[0240] The second fusion module is specifically configured to perform fusion processing on the information to be fused according to the relative position information.
[0241] Optionally, the device further includes:
[0242] A third receiving module is used to receive a second input from the user;
[0243] a second response module, configured to, in response to the second input, determine that the object corresponding to the second input is an object in the first image information, or determine that the object corresponding to the second input is an object in the second image information;
[0244] If the object corresponding to the second input is an object in the first image information, amplifying the first sound information, or reducing the second sound information;
[0245] When the object corresponding to the second input is an object in the second image information, the second sound information is amplified, or the first sound information is reduced.
[0246] In an embodiment of the present application, a second device fuses the first image information and corresponding first sound information displayed by the second device with the second image information and corresponding sound information displayed by the first device to obtain a second multimedia file, thereby achieving the fusion and superposition of images and sounds from different devices to obtain a fused multimedia file. The fused multimedia file not only has real recorded images and sounds, but also has superimposed virtual images and sounds, which enhances the interaction between the second device and the first device and makes augmented reality more realistic.
[0247] The video recording device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.
[0248] The video recording device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0249] The video recording device provided in the embodiment of the present application can achieve Figures 1 to 11 To avoid repetition, the various processes implemented by the image recording device in the method embodiment are not described here.
[0250] Alternatively, as Figure 14 As shown, an embodiment of the present application further provides an electronic device 1400, which can be the first device or the second device, including a processor 1401 and a memory 1402, and the memory 1402 stores a program or instruction that can be run on the processor 1401. When the program or instruction is executed by the processor 1401, the various steps of the above-mentioned recording method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0251] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0252] Figure 15 A hardware structure diagram of an electronic device for implementing an embodiment of the present application is provided. The electronic device may be a first device or a second device.
[0253] The electronic device 150 includes but is not limited to components such as a radio frequency unit 151 , a network module 152 , an audio output unit 153 , an input unit 154 , a sensor 155 , a display unit 156 , a user input unit 157 , an interface unit 158 , a memory 159 , and a processor 160 .
[0254] Those skilled in the art will understand that the electronic device 150 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 160 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 15 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.
[0255] Wherein, when the electronic device is a first device:
[0256] The radio frequency unit 151 is configured to obtain information to be fused, where the information to be fused includes: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information;
[0257] The processor 160 is configured to: perform fusion processing on the information to be fused to obtain a fused first multimedia file;
[0258] The display unit 156 is configured to display an image corresponding to the first multimedia file.
[0259] Optionally, the radio frequency unit 151 is further configured to:
[0260] receiving first image information sent by the second device;
[0261] Acquire the first sound information and the second sound information by one of the following methods:
[0262] Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by the microphone of the first device;
[0263] Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the second device are received;
[0264] Method 3: When the second device plays the first sound information in headphone mode, the second sound information is collected by the microphone of the first device, and the first sound information sent by the second device is received;
[0265] Method 4: When the second device plays the first sound information in headphone mode, the first sound information and the second sound information sent by the second device are received.
[0266] Optionally, the radio frequency unit 151 is further configured to:
[0267] receiving relative position information sent by the second device, where the relative position information is relative position information between the first image information and the second image information;
[0268] The processor 160 is configured to perform fusion processing on the information to be fused according to the relative position information.
[0269] Optionally, the radio frequency unit 151 is further configured to: receive a first input from a user;
[0270] The processor 160 is configured to: in response to the first input, determine that the object corresponding to the first input is an object in the first image information, or determine that the object corresponding to the first input is an object in the second image information;
[0271] If the object corresponding to the first input is an object in the first image information, amplifying the first sound information or reducing the second sound information;
[0272] When the object corresponding to the first input is an object in the second image information, the second sound information is amplified, or the first sound information is reduced.
[0273] Wherein, in the case where the electronic device is a second device:
[0274] The radio frequency unit 151 is configured to obtain information to be fused, the information to be fused including: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information;
[0275] The processor 160 is configured to perform fusion processing on the information to be fused to obtain a fused second multimedia file.
[0276] Optionally, the radio frequency unit 151 is specifically configured to:
[0277] receiving second image information sent by the first device;
[0278] Acquire the first sound information and the second sound information, or acquire the second sound information, in one of the following ways:
[0279] Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the microphone of the first device are received;
[0280] Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by a microphone of the second device;
[0281] Method three: When the second device plays the first sound information in headphone mode, receiving the second sound information collected and sent by the first device;
[0282] Method 4: When the second device plays the first sound information in headphone mode, the second sound information is collected through the microphone of the second device.
[0283] Optionally, the processor 160 is further configured to:
[0284] determining relative position information of the first image information and the second image information;
[0285] The fusing the information to be fused includes:
[0286] The information to be fused is fused according to the relative position information.
[0287] Optionally, the radio frequency unit 151 is further configured to:
[0288] receiving a second input from the user;
[0289] The processor 160 is further configured to: in response to the second input, determine that the object corresponding to the second input is an object in the first image information, or determine that the object corresponding to the second input is an object in the second image information;
[0290] If the object corresponding to the second input is an object in the first image information, amplifying the first sound information, or reducing the second sound information;
[0291] When the object corresponding to the second input is an object in the second image information, the second sound information is amplified, or the first sound information is reduced.
[0292] In an embodiment of the present application, the first device fuses the first image information and the corresponding first sound information displayed by the second device with the second image information and the corresponding sound information displayed by the first device to obtain a first multimedia file, and displays the image in the first multimedia file, thereby realizing the fusion and superposition of images and sounds of different devices. The obtained multimedia file not only has real recorded images and sounds, but also has superimposed virtual images and sounds. The first device can simultaneously synchronize the virtual content of the second device when recording audio and video, thereby strengthening the interaction between the second device and the first device, and making augmented reality more realistic.
[0293] It should be understood that in an embodiment of the present application, the input unit 154 may include a graphics processing unit (GPU) 1541 and a microphone 1542, and the graphics processor 1541 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 156 may include a display panel 1561, and the display panel 1561 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 157 includes a touch panel 1571 and other input devices 1572. The touch panel 1571 is also called a touch screen. The touch panel 1571 may include two parts: a touch detection device and a touch controller. Other input devices 1572 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
[0294] The memory 159 can be used to store software programs and various data, including but not limited to application programs and operating systems. The memory 159 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store the operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 159 may include a volatile memory or a non-volatile memory, or the memory 159 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 159 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0295] Processor 160 may include one or more processing units. Optionally, processor 160 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 160.
[0296] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned recording method embodiment is implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0297] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0298] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned recording method embodiment and achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0299] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0300] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned recording method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0301] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0302] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0303] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A recording method, applied to a first device, wherein the first device is wirelessly connected to a second device, the first device is a mobile terminal, and the second device is an augmented reality (AR) device, characterized in that: The video recording method comprises: Acquire information to be fused, where the information to be fused includes: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information; Performing fusion processing on the information to be fused to obtain a fused first multimedia file; displaying an image corresponding to the first multimedia file; The obtaining of information to be fused includes: receiving first image information sent by the second device; Acquire the first sound information and the second sound information by one of the following methods: Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by the microphone of the first device; Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the second device are received; Method 3: When the second device plays the first sound information in headphone mode, the second sound information is collected by the microphone of the first device, and the first sound information sent by the second device is received; Method 4: When the second device plays the first sound information in headphone mode, the first sound information and the second sound information sent by the second device are received.
2. The method according to claim 1, characterized in that The method further comprises: receiving relative position information sent by the second device, where the relative position information is relative position information between the first image information and the second image information; The fusing the information to be fused includes: The information to be fused is fused according to the relative position information.
3. The method according to claim 1, characterized in that The method further comprises: receiving a first input from a user; In response to the first input, determining that the object corresponding to the first input is an object in the first image information, or determining that the object corresponding to the first input is an object in the second image information; If the object corresponding to the first input is an object in the first image information, amplifying the first sound information or reducing the second sound information; When the object corresponding to the first input is an object in the second image information, the second sound information is amplified, or the first sound information is reduced.
4. A recording method, applied to a second device, the second device being wirelessly connected to a first device, the second device being an augmented reality (AR) device, and the first device being a mobile terminal, characterized in that: The video recording method comprises: Acquire information to be fused, where the information to be fused includes: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information; Performing fusion processing on the information to be fused to obtain a fused second multimedia file; The obtaining of information to be fused includes: receiving second image information sent by the first device; Acquire the first sound information and the second sound information, or acquire the second sound information, in one of the following ways: Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the microphone of the first device are received; Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by a microphone of the second device; Method three: When the second device plays the first sound information in headphone mode, receiving the second sound information collected and sent by the first device; Method 4: When the second device plays the first sound information in headphone mode, the second sound information is collected through the microphone of the second device.
5. The method according to claim 4, characterized in that The method further comprises: determining relative position information of the first image information and the second image information; The fusing the information to be fused includes: The information to be fused is fused according to the relative position information.
6. The method according to claim 4, characterized in that The method further includes: receiving a second input from a user; In response to the second input, determining that the object corresponding to the second input is an object in the first image information, or determining that the object corresponding to the second input is an object in the second image information; If the object corresponding to the second input is an object in the first image information, amplifying the first sound information, or reducing the second sound information; When the object corresponding to the second input is an object in the second image information, the second sound information is amplified, or the first sound information is reduced.
7. A video recording device, applied to a first device, the first device being wirelessly connected to a second device, the first device being a mobile terminal, and the second device being an augmented reality (AR) device, characterized in that: The video recording device comprises: A first acquisition module is configured to acquire information to be fused, the information to be fused comprising: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information; A first fusion module is used to perform fusion processing on the information to be fused to obtain a fused first multimedia file; A display module, configured to display an image corresponding to the first multimedia file; The first acquisition module is specifically configured to: receiving first image information sent by the second device; Acquire the first sound information and the second sound information by one of the following methods: Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by the microphone of the first device; Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the second device are received; Method 3: When the second device plays the first sound information in headphone mode, the second sound information is collected by the microphone of the first device, and the first sound information sent by the second device is received; Method 4: When the second device plays the first sound information in headphone mode, the first sound information and the second sound information sent by the second device are received.
8. A video recording device, applied to a second device, the second device being wirelessly connected to a first device, the second device being an augmented reality (AR) device, and the first device being a mobile terminal, characterized in that: The video recording device comprises: a second acquisition module, configured to acquire information to be fused, the information to be fused comprising: first image information displayed by the second device, first sound information corresponding to the first image information, second image information displayed by the first device, and second sound information corresponding to the second image information; A second fusion module is used to perform fusion processing on the information to be fused to obtain a fused second multimedia file; The second acquisition module is specifically used for: receiving second image information sent by the first device; Acquire the first sound information and the second sound information, or acquire the second sound information, in one of the following ways: Method 1: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information collected and sent by the microphone of the first device are received; Method 2: When the second device plays the first sound information in an external speaker mode, the first sound information and the second sound information are collected by a microphone of the second device; Method three: When the second device plays the first sound information in headphone mode, receiving the second sound information collected and sent by the first device; Method 4: When the second device plays the first sound information in headphone mode, the second sound information is collected through the microphone of the second device.
9. An electronic device, characterized in that: The present invention comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the recording method according to any one of claims 1 to 3 are implemented, or the steps of the recording method according to any one of claims 4 to 6 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the recording method according to any one of claims 1 to 3 are implemented, or the steps of the recording method according to any one of claims 4 to 6 are implemented.
Citation Information
Patent Citations
Image real-time processing method and device based on mobile terminal
CN105681684A
Karaoke system and apparatus using augmented reality, karaoke service method thereof
KR1020120081874A