A VR video processing method, device, system and storage medium

By obtaining and updating the perspective scene information in VR videos, the problem that existing VR house viewing technology cannot update the internal scenes of the house in real time is solved, and the effect of users observing the changes in spatial scenes over time in VR videos is achieved, which improves the house viewing experience.

CN117292090BActive Publication Date: 2025-05-30BANGLIDE TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311307285.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-10
Publication Date
2025-05-30
Estimated Expiration
2043-10-10

AI Technical Summary

Technical Problem

The existing VR house viewing technology cannot update the interior scenes of the house in real time, resulting in the inability of viewers to observe the scenes changing within the house over time, affecting the house viewing experience.

Method used

Through a VR video processing method, the second video information or image information associated with the multiple viewing angle scenes of the initial VR video are obtained, and the first VR video is updated, so that the user can see the scene transitioning from the scene of the initial acquisition time to the second acquisition time when watching.

Benefits of technology

It realizes the effect of users observing the changes in space scenes over time in VR videos, improves the intuitiveness and time extension of the space, and improves the house viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117292090B_ABST
    Figure CN117292090B_ABST
Patent Text Reader

Abstract

The VR video processing method of the present invention realizes the update of multiple perspective scenes of the first VR video by obtaining second video information or image information associated with multiple perspective scenes of the first VR video, so as to achieve that when the user stays at a certain perspective during the process of watching the first VR video, the scene from the first acquisition time is shown to transition to the scene at the second acquisition time, so that the user can obtain the change process of the spatial scene content over time based on the timeline, improving the intuitiveness and time extensibility of the user's observation of the space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of VR video processing, and in particular, to a VR video processing method, apparatus, system, and storage medium. Background Art

[0002] Virtual reality technology (English name: Virtual Reality, abbreviated as VR), also known as virtual reality or immersive technology, is mainly realized by computer technology, making use of and integrating the latest development achievements of various high-techs such as 3D graphics technology, multimedia technology, simulation technology, and display technology. With the help of devices such as computers, a virtual world with a realistic three-dimensional visual, tactile, olfactory, and other sensory experiences is generated, so that people in the virtual world have a sense of being on the spot.

[0003] In the prior art, VR technology has been widely applied to the real estate field. For example, VR house viewing allows users to freely move and observe each room and see the real scene of the room in all directions through VR virtual reality technology and 3D panoramic display technology. VR house viewing is a new technology in the real estate industry. House viewers can view various information about the house through a mobile phone or a tablet computer, see what the overall internal structure of the house is like, and do not need to go to the site to view, which is very time-saving, simple, and fast. However, in existing VR house viewing, real estate agents pre-record 3D images of houses through professional recording equipment and upload them to housing transaction intermediary platforms such as Beike. House viewers can click on the VR house viewing function of houses on platforms such as Beike to achieve VR house viewing. As the housing transaction cycle lengthens, for example, ranging from several months to several years, some of the captured internal scenes of the house may have changed significantly with the extension of the residence time. As a result, house viewers cannot observe the changed scenes inside the house through the VR house viewing function in the later stage, seriously affecting the house viewing experience of house viewers.

[0004] Therefore, there is a need to propose a VR video processing method to solve the above problems. Summary of the Invention

[0005] The present invention provides a VR video processing method, and the method includes the following steps:

[0006] Step S1, a first user collects spatial scene information of a first space to generate a first VR video, and the first VR video includes first identification information and first acquisition time information of the first space;

[0007] Step S2, a second user collects video information or image information of a second space, and the video information or image information includes second identification information and second acquisition time information of the second space;

[0008] In step S3, determine whether the second identification information is the same as the first identification information. If they are the same, then execute step S4; otherwise, save the video information or the image information.

[0009] In step S4, determine whether the second acquisition time information is later than the first acquisition time information. If it is, then execute step S5; otherwise, end the process.

[0010] In step S5, update the first VR video with the video information or the image information, so that when the third user stays at a certain perspective scene during the viewing of the first VR video, the scene from the first acquisition time is shown to transition to the scene at the second acquisition time.

[0011] As a preferred implementation manner, after determining whether the second identification information is the same as the first identification information, if they are the same, it further includes:

[0012] Establish a timeline based on the first acquisition time information and the second acquisition time information;

[0013] Establish an association relationship between the first VR video and the timeline;

[0014] Establish an association relationship between the video information or the image information and the timeline.

[0015] As a preferred implementation manner, when updating the first VR video with the video information, it further includes:

[0016] Obtain multiple key perspective information in the first VR video, and obtain the corresponding perspective scenes in the first VR video according to the multiple key perspective information;

[0017] Search for perspective information in the video information that matches the key perspective information, and establish a corresponding relationship between each key perspective information and the matching perspective information;

[0018] Update the corresponding perspective scenes in the first VR video with the perspective scenes associated with the matching perspective information according to the corresponding relationship.

[0019] As a preferred implementation manner, when searching for perspective information in the video information that matches the key perspective information and establishing a corresponding relationship between each key perspective information and the matching perspective information, it further includes:

[0020] Mark the key perspective information at the position of the first acquisition time on the timeline to obtain multiple key perspective marks;

[0021] Search for perspective information that matches the key perspective information from the video information, and mark the matching perspective information at the second acquisition time position of the timeline to obtain a plurality of matching perspective identifiers;

[0022] Establish the corresponding relationship between each of the matching perspective identifiers and the corresponding key perspective identifier.

[0023] As a preferred embodiment, using the image information to update the first VR video further includes:

[0024] Obtain a plurality of key video frames in the first VR video, and search for image information that matches the key video frames from the images;

[0025] Use the matching images to update the corresponding key video frames in the first VR video.

[0026] As a preferred embodiment, using the matching images to update the corresponding key video frames in the first VR video further includes:

[0027] Mark the key video frames at the first acquisition time position of the timeline to obtain a plurality of key video frame identifiers;

[0028] Search for image information that matches the key video frame information from the image information, and mark the matching image information at the second acquisition time position of the timeline to obtain a plurality of matching image identifiers;

[0029] Establish the corresponding relationship between each of the matching image identifiers and the corresponding key image identifier.

[0030] As another embodiment, the present invention provides a VR video processing device, and the device includes the following modules:

[0031] A VR video acquisition module, configured to acquire spatial scene information of a first space by a first user to generate a first VR video, where the first VR video includes first identification information and first acquisition time information of the first space;

[0032] An updated data acquisition module, configured to acquire video information or image information of a second space by a second user, where the video information or image information includes second identification information and second acquisition time information of the second space;

[0033] A first determination module, configured to determine whether the second identification information is the same as the first identification information. If they are the same, execute the second determination module; otherwise, save the video information or the image information;

[0034] A second judgment module, configured to judge whether the second acquisition time information is later than the first acquisition time information. If so, execute the VR video update module;

[0035] A VR video update module, which updates the first VR video using the video information or the image information, so that when the third user stays at a certain perspective scene during the process of watching the first VR video, the scene from the first acquisition time is displayed to transition to the scene at the second acquisition time.

[0036] As a preferred implementation manner, the first judgment module further includes:

[0037] Establish a timeline based on the first acquisition time information and the second acquisition time information;

[0038] Establish an association relationship between the first VR video and the timeline;

[0039] Establish an association relationship between the video information or the image information and the timeline.

[0040] As a preferred implementation manner, when using the video information to update the first VR video, it further includes:

[0041] Obtain multiple key perspective information in the first VR video, and obtain the corresponding perspective scenes in the first VR video according to the multiple key perspective information;

[0042] Search for perspective information matching the key perspective information from the video information, and establish a corresponding relationship between each key perspective information and the matching perspective information;

[0043] Update the corresponding perspective scenes in the first VR video with the perspective scenes associated with the matching perspective information according to the corresponding relationship.

[0044] As a preferred implementation manner, when searching for perspective information matching the key perspective information from the video information and establishing a corresponding relationship between each key perspective information and the matching perspective information, it further includes:

[0045] Mark the key perspective information at the position of the first acquisition time on the timeline to obtain multiple key perspective marks;

[0046] Search for perspective information matching the key perspective information from the video information, and mark the matching perspective information at the position of the second acquisition time on the timeline to obtain multiple matching perspective marks;

[0047] Establish a corresponding relationship between each matching perspective mark and the corresponding key perspective mark.

[0048] As a preferred embodiment, updating the first VR video using the image information further includes:

[0049] Obtain multiple key video frames in the first VR video, and find image information in the image that matches the key video frames;

[0050] Update the corresponding key video frames in the first VR video using the matching images.

[0051] As a preferred embodiment, updating the corresponding key video frames in the first VR video using the matching images further includes:

[0052] Mark the key video frames at the first acquisition time position of the timeline to obtain multiple key video frame identifiers;

[0053] Find image information in the image information that matches the key video frame information, and mark the matching image information at the second acquisition time position of the timeline to obtain multiple matching image identifiers;

[0054] Establish a corresponding relationship between each of the matching image identifiers and the corresponding key image identifiers.

[0055] As another embodiment, the present invention provides a VR video processing system, and the system executes the VR video processing method described above.

[0056] As another embodiment, the present invention provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and the computer program executes the VR video processing method.

[0057] It can be seen that the above VR video processing method of the present invention realizes the update of multiple perspective scenes of the first VR video by obtaining second video information or image information associated with multiple perspective scenes of the first VR video, so as to achieve that when the user stays at a certain perspective during the process of watching the first VR video, the scene from the first acquisition time is displayed to transition to one or more scenes at the second acquisition time, for the user to obtain the change process of the spatial scene content over time based on the timeline, improving the intuitiveness and time extensibility of the user's observation of the space. Description of the Drawings

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments and the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0059] Figure 1 It is a schematic diagram of the steps of a VR video processing method of the present invention.

[0060] Figure 2 It is a schematic diagram of the timeline of the present invention and its video or image identifier.

[0061] Figure 3 It is a schematic diagram of the structure of a VR video processing device of the present invention. Detailed implementation manners

[0062] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0063] Embodiment 1:

[0064] As Figure 1 shown, the present invention provides a VR video processing method, and the method includes the following steps:

[0065] Step S1: The first user collects the spatial scene information of the first space to generate a first VR video, where the first VR video includes the first identification information and the first acquisition time information of the first space. It should be noted that the first space can be a complete housing unit. For example, it can be a small building and its internal space, or the internal space of a housing unit recorded in the property ownership certificate, without limitation here. The spatial scene is a part of the internal space. For example, if the first space is the housing space of a three-bedroom and one-living-room house, the spatial scene can be the scene information captured in any one room. The first VR video is a three-dimensional space video obtained by shooting the internal space of the building or the housing unit, and the first VR video is the video obtained by shooting the internal space of the building or the housing unit for the first time. Preferably, the first user is a staff member assigned by a housing agency, and a professional VR camera device is used to shoot the internal space of the building or the housing unit. Since the first VR video is shot by a professional staff member, the shooting quality of the first VR video is good, laying a good foundation for subsequent editing and processing based on the first VR video. Or, the first user can also be other person who shoots the internal space to generate the first VR video, without limitation here. Further, for subsequent processing, in the present invention, the first identification information and the first acquisition time information of the first space are established for the first VR video when generating the first VR video. The first identification information is used to identify the identity information of the first space, which is information that is unique and distinguishable from other spaces. For example, the property number information, property address, floor, and household number of the property ownership certificate of the building or the housing unit, or information defined by the user to identify the identity of the first space, such as a house ID, without limitation here. The first acquisition time information can be the date information of the first shooting, or the date information of generating the first VR video.

[0066] Step S2, the second user collects video information or image information of the second space, and the video information or image information includes the second identification information and the second collection time information of the second space; it should be noted that, since the first VR video is shot by a professional cameraman, the subsequent video update can be carried out on this basis, so the quality of the video information or image information required to update the first VR video does not need to be too high, so preferably, the second user is the video or image information collected by the real estate agent in the subsequent viewing process in the second space, because the real estate agent usually takes a video or photo of the house during the viewing process, and then publishes it to self-media platforms such as Douyin and WeChat to attract users. At this time, the real estate agent only needs to upload the video or photo he took to a second-hand housing transaction platform such as Beike to update the first VR video, thereby basically not increasing the workload of the real estate agent. In order to construct the first VR video based on time update, it is necessary to establish a space identification and time information for the video information or image information used for updating, and the identification method of the second identification information is the same as the identification method of the first identification information, which is not repeated here; the identification method of the second collection time information is the same as the identification method of the first collection time information, which is not repeated here. Furthermore, there may be multiple second users, so that each second user can collect corresponding video or image information, thereby updating the first VR video in a time series.

[0067] Step S3, determine whether the second identification information is the same as the first identification information. If they are the same, execute step S4, otherwise, save the video information or the image information; it should be noted that in order to avoid the video information or image information used for updating not matching the first VR video, it is necessary to perform the identification information matching process before updating, and only perform the subsequent process when the second identification information is the same as the first identification information. Otherwise, save the video information or the image information for further analysis or matching.

[0068] Step S4: Determine whether the second acquisition time information is later than the first acquisition time information. If so, execute Step S5; otherwise, end the process. It should be noted that, generally, since the first acquisition time is the time of the first shot of the internal space, and the second acquisition time is the time of collecting videos or images during subsequent property viewings, it is necessary to determine whether the second acquisition time information is later than the first acquisition time information to identify whether there is a time dislocation. When the second acquisition time is at the first acquisition time, it indicates that there is no time dislocation, and thus the subsequent process can continue. On the contrary, there is a problem of time dislocation, and at this time, further judgment is required, and thus the entire process ends. Further, if there are multiple second acquisition times, that is, there are video or image information for updating at multiple time points, then each second acquisition time is compared with the first acquisition time respectively to identify the time dislocation at that time.

[0069] Step S5: Use the video information or the image information to update the first VR video, so that when the third user stays at a certain perspective scene during the viewing of the first VR video, the scene from the first acquisition time to the second acquisition time is displayed. It should be noted that after obtaining the video or the image for updating, the first video can be updated; it should be emphasized that since the present invention is committed to presenting the effect of the spatial scene changing over time for users in various spaces, the update method of the first VR video in the present invention is not the conventional deletion and overwrite update, but the superimposed update in the same spatial scene, that is, the dynamic evolution update based on the spatial scene; thus, the third user, such as an ordinary viewer on the housing trading platform, when staying at a certain spatial scene or a certain perspective scene of a spatial scene during the viewing of the first VR video, since the scene or the perspective scene is a series of scenes related to time, the change process of the spatial scene or the perspective scene over time can be shown to the user, thereby realizing the update of the first VR video of the present invention. Among them, the superimposed update in the same spatial scene, that is, the dynamic evolution update, can be: identifying the original video and the video for updating in the same spatial scene in the order of shooting or uploading time. Thus, when the user operates the first VR video to switch to the current spatial scene, first play the original video corresponding to the current scene, and then play the videos for updating the current scene in sequence according to the time order. Among them, the original video is the video segment or sub-video corresponding to each spatial scene parsed from the first VR video.

[0070] It can be seen that the above VR video processing method of the present invention updates multiple perspective scenes of the first VR video by obtaining second video information or image information associated with the multiple perspective scenes of the first VR video, so as to achieve that when the user stays at a certain perspective during the process of watching the first VR video, the scene from the first acquisition time is displayed to transition to the scene at the second acquisition time, so that the user can obtain the change process of the spatial scene content over time based on the timeline, improving the intuitiveness and time extensibility of the user's observation of the space.

[0071] As a preferred implementation manner, it is determined whether the second identification information and the first identification information are the same. If they are the same, then, it further includes:

[0072] A timeline is established according to the first acquisition time information and the second acquisition time information; it should be noted that, for the convenience of managing the relevant information of the above video or image changing over time, preferably, the invention identifies this information through a timeline, such as Figure 2 shown, wherein, the timeline has multiple time points, and each of the time points corresponds to the first acquisition time or the second acquisition time; for example, the first acquisition time is T1, and the two second acquisition times are T2 and T3 respectively.

[0073] An association relationship between the first VR video and the timeline is established; an association relationship between the video information or the image information and the timeline is established. After multiple time points are marked on the timeline, an association relationship between each time point and the video or image needs to be established; such as Figure 2As shown, the first VR video V1 is associated with the first acquisition time T1. The first VR video V1 includes video information of three spatial scenes V11, V12, and V13. The second video V2 is associated with the second acquisition time T2. The second video V2 includes video information of three spatial scenes V21, V22, and V23. The image set V3 is associated with the third acquisition time T3. The image set V3 includes image subset information of three spatial scenes V31, V32, and V33. Further, the above association relationships can be respectively represented by the following data structures: [T1, V1, V11-V12-V13], where the time point on the timeline is represented by the first digit T1, the second digit V1 represents the video set identifier, and the third digit V11-V12-V13 are multiple spatial scene identifiers in the video and their corresponding video subsets; [T2, V2, V21-V22-V23], where the time point on the timeline is represented by the first digit T2, the second digit V2 represents the video set identifier, and the third digit V21-V22-V23 are multiple spatial scene identifiers in the video and their corresponding video subsets; [T3, V3, V31-V32-V33], where the time point on the timeline is represented by the first digit T3, the second digit V3 represents the image set identifier, and the third digit V11-V12-V13 are multiple spatial scene identifiers in the image set and their corresponding image subsets.

[0074] As a preferred implementation manner, updating the first VR video using the video information further includes:

[0075] Obtain multiple key perspective information in the first VR video, and obtain the corresponding perspective scene in the first VR video according to the multiple key perspective information. It should be noted that since the video viewer can switch perspectives in a certain scene of the VR video through operations such as sliding, it is necessary to obtain the perspective information in the scene. And different perspectives present different amounts of information to the user, and the importance levels are also different. Therefore, the present invention updates a certain scene by obtaining the perspective scene of the key perspective. For example, Figure 2As shown, the identifier Vx1 is used to represent the living room space, where x represents the time point identifier. The key perspective is the perspective in a certain space (such as the living room) that can maximize the overview of important or most item information. Specifically, all video frames of the living room captured in the first VR video are obtained, and these video frames are compared to filter out the video frames with the most items captured from these video frames, or to filter out the video frames that can simultaneously see important items such as sofas, coffee tables, TV cabinets, and chandeliers. Thus, one or more perspectives corresponding to the filtered video frames are used as the key perspective information of the living room, and the perspective scenes corresponding to these perspectives in the first VR video are obtained, so as to facilitate the user to observe the layout and items of the living room as a whole, achieving the purpose of updating the first VR video with the least amount of video or image information.

[0076] Search for the perspective information that matches the key perspective information from the video information, and establish the corresponding relationship between each key perspective information and the matching perspective information. At this time, since the first VR video is updated using the video information, the videos used for updating (such as Figure 2Find the perspective information that matches the key perspective information in the second video V2 shown at the T2 time point, and establish the corresponding relationship between each piece of the key perspective information and the matching perspective information. For example, continuing with the foregoing embodiment, the two perspectives corresponding to the filtered video frames are used as the key perspective information of the living room (assuming the identifier of the spatial scene of the living room is V11). Thus, adding one digit after the identifier V11 gives V111 and V112 to represent the two key perspective information in the spatial scene of the V11 living room. For example, the identifier V111 represents the perspective of observing the living room from the balcony and being able to see important items such as the sofa, coffee table, TV cabinet, and chandelier, and the identifier V112 represents the perspective from the opposite side of the balcony where the most items in the living room can be seen. And a similar method can be used to obtain the key perspective information of the second video, and by adding one digit after the identifier V21 of the spatial scene of the living room in the second video, namely V211 and V212, to represent the two key perspective information in the spatial scene of the V21 living room. For example, the identifier V211 represents the perspective of observing the living room from the balcony and being able to see important items such as the sofa, coffee table, TV cabinet, and chandelier, and the identifier V212 represents the perspective from the opposite side of the balcony where the most items in the living room can be seen. It can be seen that the identifier V111 and the identifier V211 have the same identifier digit Vx11, so as to be able to simultaneously identify the perspective of observing the living room from the balcony; the identifier V112 and the identifier V212 have the same identifier digit Vx12, so as to be able to simultaneously identify the perspective of observing the living room from the opposite side of the balcony. Thus, through the last identifier digit above, the corresponding relationship between the key perspective information and the matching perspective information can be established. Then, update the corresponding perspective scene in the first VR video with the perspective scene associated with the matching perspective information according to the corresponding relationship. The specific update method is the same as that in the foregoing embodiment and will not be elaborated here.

[0077] As a preferred implementation manner, finding the perspective information that matches the key perspective information from the video information and establishing the corresponding relationship between each piece of the key perspective information and the matching perspective information further includes:

[0078] Mark the key perspective information at the first acquisition time position of the timeline to obtain a plurality of key perspective identifiers. It should be noted that mark the key perspective information at the first acquisition time position T1 of the timeline to obtain a plurality of key perspective identifiers. For example, use V111 and V112 to represent the two key perspective identifier information in the spatial scene of the V11 living room in the first VR video.

[0079] Find the perspective information that matches the key perspective information from the video information, and mark the matching perspective information at the second acquisition time position of the timeline to obtain multiple matching perspective identifiers. It should be noted that after establishing the key perspective identifier, it is necessary to find the key perspective of the corresponding spatial scene from the second video for updating. The specific search method can be determined by comparing key video frames. For example, first determine the key video frames corresponding to multiple key videos of each spatial scene in the first VR video, and then compare these key frames with the video frames of the second video. When the similarity of the video frames is greater than the set threshold, it is determined that the match is successful; otherwise, the match fails. Mark the key video frames that match successfully in the second video frame. For example, mark the key video frames that match successfully in the second video as V211 and V212, so as to establish the corresponding relationship between each matching perspective identifier and the corresponding key perspective identifier. Since each spatial scene in the first VR video already has multiple key perspectives, the key video frames in the second video also have the corresponding key perspectives and their spatial scenes, so they can be directly marked. Further, considering that the items in the room gradually change over time, in order to improve the matching efficiency, the video corresponding to the key video frames used to match the updated video can be the video that is closest to the current time point and earlier than the current time point on the timeline.

[0080] As a preferred implementation manner, using the image information to update the first VR video further includes:

[0081] Obtain multiple key video frames in the first VR video, and find the image information that matches the key video frames from the image. It should be noted that in this embodiment, the first VR video is updated using image information because the intermediary who will show the house to the customer later may not be good at or convenient to shoot house videos, thus providing another implementation manner. Since the information that can be provided by the image is less, this embodiment no longer updates the first VR video from the perspective, but only updates it based on the comparison between the key video frames and the image. Specifically, obtain multiple key video frames in the first VR video. For example, obtain one or more key video frames in each spatial scene respectively, and then find the image information that matches the key video frames from the image. Among them, the determination method of the key video frames is the same as that in the previous embodiment and will not be elaborated here. The image can be determined whether it matches the key video frames by setting a similarity threshold, so as to update the corresponding key video frames in the first VR video with the matching image. Further, in order to reduce the jerks of the VR video caused by the delayed-playing images superimposed on the original key video frames, the VR fusion or image fusion method can be used to improve the visual effect of the above update.

[0082] As a preferred embodiment, updating the corresponding key video frames in the first VR video using matching images further includes:

[0083] Identifying the key video frames at the first acquisition time position of the timeline to obtain a plurality of key video frame identifiers; it should be noted that the image-based update method can manage the updated images using a timeline and its identification method similar to that of the video-based update method. Among them, the identification method of the first VR video on the timeline is the same as that in the foregoing embodiment. Identifying the key video frames at the first acquisition time position of the timeline to obtain a plurality of key video frame identifiers will not be elaborated here.

[0084] Searching for image information matching the key video frame information from the image information, identifying the matching image information at the second acquisition time position of the timeline to obtain a plurality of matching image identifiers; establishing a corresponding relationship between each of the matching image identifiers and the corresponding key image identifier. It should be noted that when the image for update successfully matches the key video frame, the identification information of the corresponding successfully matched image is established according to the identification information of the key video frame. For example, as Figure 2 shown, the third acquisition time T3 is associated with an image set V3, and the image set V3 includes image subset information of three spatial scene images V31, V32, and V33. Further, considering that the items in the room gradually change over time, in order to improve the matching efficiency, the video corresponding to the key video frame used for matching with the updated image can be the video that is closest to the current time point and earlier than the current time point on the timeline.

[0085] It can be seen that the above VR video processing method of the present invention realizes the update of multiple perspective scenes of the first VR video by obtaining second video information or image information associated with multiple perspective scenes of the first VR video, so as to achieve the display of the transition from the scene at the first acquisition time to one or more scenes at the second acquisition time when the user stays at a certain perspective during the process of watching the first VR video, for the user to obtain the change process of the spatial scene content over time based on the timeline, improving the intuitiveness and time extensibility of the user's observation of the space.

[0086] Embodiment 2:

[0087] As Figure 3 shown, the present invention provides a VR video processing device, and the device includes the following modules:

[0088] A VR video acquisition module is used for a first user to collect spatial scene information of a first space to generate a first VR video, and the first VR video includes first identification information and first acquisition time information of the first space. It should be noted that the first space can be a complete housing unit. For example, it can be a small building and its internal space, or the internal space of a housing unit recorded in a property ownership certificate, without limitation here. The spatial scene is a part of the internal space. For example, if the first space is the housing space of a three-bedroom and one-living-room house, the spatial scene can be the scene information captured in any one room. The first VR video is a three-dimensional space video obtained by shooting the internal space of the building or housing unit, and the first VR video is the video obtained by shooting the internal space of the building or housing unit for the first time. Preferably, the first user is a staff member assigned by a housing agency, so as to use professional VR camera equipment to shoot the internal space of the building or housing unit. Since the first VR video is shot by professional staff, the shooting quality of the first VR video is good, laying a good foundation for subsequent editing and processing based on the first VR video. Or, the first user is other personnel who shoot the internal space for the first time to generate the first VR video, without limitation here. Further, for subsequent further processing, the present invention establishes the first identification information and the first acquisition time information of the first space for the first VR video when generating the first VR video. The first identification information is information used to identify the identity of the first space, such as the property number information of the property ownership certificate of the building or the housing unit, or information defined by the user to identify the identity of the first space, such as ID, without limitation here. The first acquisition time information can be the date information of the first shooting, or the date information of generating the first VR video.

[0089] Update data acquisition module, used for the second user to collect video information or image information of the second space, the video information or image information includes the second identification information and the second collection time information of the second space; it should be noted that, since the first VR video is shot by a professional cameraman, the subsequent video update can be carried out on this basis, so the quality of the video information or image information required to update the first VR video does not need to be too high, so preferably, the second user is the video or image information collected by the real estate agent in the subsequent viewing process in the second space, because the house is usually video-shot or photographed during the viewing process, and then published to self-media platforms such as Douyin and WeChat to attract users. At this time, the real estate agent only needs to upload the video or photo he took to the second-hand housing transaction platform such as Beike to realize the update of the first VR video, thereby basically not increasing the workload of the real estate agent. In order to construct the first VR video based on time update, it is necessary to establish a space identification and time information for the video information or image information used for updating, and the identification method of the second identification information is the same as the identification method of the first identification information, which is not repeated here; the identification method of the second collection time information is the same as the identification method of the first collection time information, which is not repeated here. Furthermore, there may be multiple second users, so that each second user can collect corresponding video or image information, thereby updating the first VR video in a time series.

[0090] The first judgment module is used to judge whether the second identification information is the same as the first identification information. If they are the same, the second judgment module is executed. Otherwise, the video information or the image information is saved. It should be noted that in order to avoid the video information or image information used for updating not matching the first VR video, it is necessary to perform a matching process of the identification information before updating. The subsequent process is only performed when the second identification information is the same as the first identification information. Otherwise, the video information or the image information is saved for further analysis or matching.

[0091] A second judgment module, configured to judge whether the second acquisition time information is later than the first acquisition time information. If so, the VR video update module is executed. It should be noted that, usually, since the first acquisition time is the time when the internal space is first photographed, and the second acquisition time is the time when videos or images are acquired during subsequent property viewings, it is necessary to judge whether the second acquisition time information is later than the first acquisition time information to identify whether there is a time dislocation. When the second acquisition time is the same as the first acquisition time, it indicates that there is no time dislocation, and the subsequent process can continue. On the contrary, there is a problem of time dislocation. At this time, further judgment is required to end the entire process. Further, if there are multiple second acquisition times, that is, there are video or image information for updating at multiple time points, each second acquisition time is respectively compared with the first acquisition time to identify the time dislocation at that time.

[0092] A VR video update module, which uses the video information or the image information to update the first VR video, so that when the third user stays at a certain perspective scene during the viewing of the first VR video, the scene from the first acquisition time to the second acquisition time is displayed. It should be noted that after obtaining the video or the image for updating, the first video can be updated; it should be emphasized that since the present invention is committed to presenting the effect of the spatial scene changing over time for users in various spaces, the update method of the first VR video in the present invention is not the conventional deletion and overwrite update, but the overlay update in the same spatial scene, that is, the dynamic evolution update based on the spatial scene; thus, the third user, such as an ordinary viewer on the housing transaction platform, when staying at a certain spatial scene or a certain perspective scene of a spatial field during the viewing of the first VR video, since the scene or the perspective scene is a series of scenes related to time, the change process of the spatial scene or the perspective scene over time can be shown to the user, thereby realizing the update of the first VR video of the present invention.

[0093] It can be seen that the above VR video processing method of the present invention realizes the update of multiple perspective scenes of the first VR video by obtaining the second video information or image information associated with multiple perspective scenes of the first VR video, so as to achieve that when the user stays at a certain perspective during the viewing of the first VR video, the scene from the first acquisition time to the second acquisition time is displayed, so that the user can obtain the change process of the spatial scene content over time based on the time line, improving the intuitiveness and time extensibility of the user's observation of the space.

[0094] As a preferred embodiment, it is determined whether the second identification information is the same as the first identification information. If they are the same, then, further comprising:

[0095] A timeline is established according to the first acquisition time information and the second acquisition time information. It should be noted that, for the convenience of managing the relevant information of the above-mentioned video or image changing with time, preferably, the invention identifies this information through a timeline, such as Figure 2 shown, wherein the timeline has multiple time points, and each time point corresponds to one of the first acquisition time or the second acquisition time; for example, the first acquisition time is T1, and the two second acquisition times are T2 and T3 respectively.

[0096] An association relationship between the first VR video and the timeline is established; an association relationship between the video information or the image information and the timeline is established. After multiple time points are marked on the timeline, an association relationship between each time point and the video or image needs to be established; such as Figure 2 shown, the first VR video V1 is associated at the first acquisition time T1, and the first VR video V1 includes video information of three spatial scenes V11, V12, and V13; the second video V2 is associated at the second acquisition time T2, and the second video V2 includes video information of three spatial scenes V21, V22, and V23; the image set V3 is associated at the third acquisition time T3, and the image set V3 includes image subset information of three spatial scenes V31, V32, and V33. Further, the above association relationship can be represented by the following fields: [T1, V1, V11-V12-V13], where the time point on the timeline is represented by the first digit T1, the second digit V1 represents the video or image set identifier, and the third digit V11-V12-V13 are the identifiers of multiple spatial scenes in the video or image set and their corresponding video or image subsets.

[0097] As a preferred embodiment, when using the video information to update the first VR video, it further comprises:

[0098] Multiple key perspective information in the first VR video is obtained, and the corresponding perspective scene in the first VR video is obtained according to the multiple key perspective information. It should be noted that since a video viewer can switch perspectives in a certain scene of a VR video by sliding or other operations, it is necessary to obtain the perspective information in the scene. And different perspectives present different amounts of information to the user, and the importance is also different. Therefore, the invention updates a certain scene by obtaining the perspective scene of the key perspective. For example, such as Figure 2As shown, the identifier Vx1 is used to represent the living room space, where x represents the time point identifier; then the key perspective is the perspective that can maximize the overview of important or most item information. For example, one or more perspectives that can simultaneously see the sofa, coffee table, TV cabinet, and chandelier, and based on these perspectives, obtain the corresponding perspective scenes in the first VR video, so as to facilitate the user to observe the layout and items of the living room as a whole, achieving the purpose of updating the first VR video with the least amount of video or image information.

[0099] Search for perspective information that matches the key perspective information from the video information, and establish the corresponding relationship between each key perspective information and the matching perspective information; at this time, since the video information is used to update the first VR video, these videos used for updating (such as Figure 2 the second video V2 at time point T2 shown) are used to search for perspective information that matches the key perspective information, and establish the corresponding relationship between each key perspective information and the matching perspective information; for example, add one digit V111, V112 after the identifier V11 to represent two key perspective information in the V11 living room space scene, and correspondingly, add one digit V211, V212 after the identifier V21 in the second video to represent two key perspective information in the V21 living room space scene. The corresponding relationship between the key perspective information and the matching perspective information can be established through the last digit identifier. Then, update the corresponding perspective scene in the first VR video with the perspective scene associated with the matching perspective information according to the corresponding relationship. The specific update method is the same as that in the foregoing embodiment and will not be elaborated here.

[0100] As a preferred implementation manner, searching for perspective information that matches the key perspective information from the video information and establishing the corresponding relationship between each key perspective information and the matching perspective information further includes:

[0101] Mark the key perspective information at the first acquisition time position of the time line to obtain a plurality of key perspective identifiers; it should be noted that mark the key perspective information at the first acquisition time position T1 of the time line to obtain a plurality of key perspective identifiers. For example, use V111, V112 to represent two key perspective identifier information in the V11 living room space scene in the first VR video.

[0102] Find the perspective information that matches the key perspective information from the video information, and identify the matching perspective information at the second acquisition time position of the timeline to obtain multiple matching perspective identifiers. It should be noted that after establishing the key perspective identifier, it is necessary to find the key perspective of the corresponding spatial scene from the second video for updating. The specific search method can be determined by comparing key video frames. For example, first determine the key video frames corresponding to multiple key videos of each spatial scene in the first VR video, and then compare these key frames with the video frames of the second video. When the similarity of the video frames is greater than the set threshold, it is determined that the match is successful; otherwise, the match fails. Identify the key video frames that match successfully in the second video frame. For example, identify the key video frames that match successfully in the second video as V211 and V212, so as to establish the corresponding relationship between each matching perspective identifier and the corresponding key perspective identifier. Since each spatial scene in the first VR video already has multiple key perspectives, the key video frames in the second video also have the corresponding key perspectives and their spatial scenes, so they can be directly identified. Further, considering that the items in the room gradually change over time, in order to improve the matching efficiency, the video corresponding to the key video frame used to match the updated video can be the video that is closest to the current time point and earlier than the current time point on the timeline.

[0103] As a preferred implementation manner, using the image information to update the first VR video further includes:

[0104] Obtain multiple key video frames in the first VR video, and find the image information that matches the key video frames from the image. It should be noted that in this embodiment, the first VR video is updated using image information because the intermediary who will show the house to the customer later may not be good at or inconvenient to shoot house videos, thus providing another implementation manner. Since the information that can be provided by images is less, this embodiment no longer updates the first VR video from the perspective angle, but only updates it based on the comparison between key video frames and images. Specifically, obtain multiple key video frames in the first VR video. For example, obtain one or more key video frames in each spatial scene respectively, and then find the image information that matches the key video frames from the image. Among them, the method for determining the key video frames is the same as that in the previous embodiment and will not be elaborated here. The similarity threshold can be set to determine whether the image matches the key video frame, so as to update the corresponding key video frame in the first VR video with the matching image. Further, in order to reduce the jerks of the VR video caused by the images with delayed playback superimposed on the original key video frames, the VR fusion or image fusion method can be used to improve the updated visual effect.

[0105] As a preferred embodiment, updating the corresponding key video frames in the first VR video using the matching images further includes:

[0106] Identifying the key video frames at the first acquisition time position of the timeline to obtain a plurality of key video frame identifications; it should be noted that the image-based update method can be managed for updating images using a timeline and its identification method similar to that of the video-based update method. Among them, the identification method of the first VR video on the timeline is the same as that in the foregoing embodiment. Identifying the key video frames at the first acquisition time position of the timeline to obtain a plurality of key video frame identifications will not be elaborated here.

[0107] Searching for the image information matching the key video frame information from the image information, identifying the matching image information at the second acquisition time position of the timeline to obtain a plurality of matching image identifications; establishing a corresponding relationship between each of the matching image identifications and the corresponding key image identification. It should be noted that after the image for update successfully matches the key video frame, the identification information of the successfully matched image is established according to the identification information of the key video frame. For example, as Figure 2 shown, the image set V3 is associated at the third acquisition time T3, and the image set V3 includes the image subset information of three spatial scene images V31, V32, and V33. Further, considering that the items in the room gradually change over time, in order to improve the matching efficiency, the video corresponding to the key video frame used for matching with the updated image can be the video that is closest to the current time point and earlier than the current time point on the timeline.

[0108] It can be seen that the above VR video processing device of the present invention realizes the update of multiple perspective scenes of the first VR video by obtaining the second video information or image information associated with multiple perspective scenes of the first VR video, so as to achieve the transition of the scene from the first acquisition time to one or more of the second acquisition times when the user stays at a certain perspective during the process of watching the first VR video, for the user to obtain the change process of the spatial scene content over time based on the timeline, improving the intuitiveness and time extensibility of the user's observation of the space.

[0109] Embodiment 3:

[0110] As another embodiment, the present invention provides a VR video processing system, and the system executes the VR video processing method of the foregoing embodiment.

[0111] Those skilled in the art can understand that the present invention includes devices for performing one or more of the operations described in this application. These devices can be specially designed and manufactured for the required purposes, or they can also include known devices in general-purpose computers. These devices have computer programs stored therein, and these computer programs are selectively activated or reconstructed. Such computer programs can be stored in a device (e.g., a computer) readable medium or in any type of medium suitable for storing electronic instructions and coupled to the bus respectively. The computer readable medium includes, but is not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. That is, the readable medium includes any medium that stores or transmits information in a form readable by a device (e.g., a computer).

[0112] Those skilled in the art can understand that each block in these structural diagrams and / or block diagrams and / or flowcharts, as well as combinations of blocks in these structural diagrams and / or block diagrams and / or flowcharts, can be implemented with computer program instructions. Those skilled in the art can understand that these computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing methods to be implemented, so that the solutions specified in the blocks or multiple blocks of the structural diagrams and / or block diagrams and / or flowcharts disclosed by the present invention are executed by the processor of the computer or other programmable data processing methods.

[0113] Those skilled in the art can understand that the various operations, methods, steps, measures, and solutions in the processes discussed in the present invention can be alternated, changed, combined, or deleted. Further, other steps, measures, and solutions in the various operations, methods, and processes discussed in the present invention can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those disclosed in the various operations, methods, and processes of the present invention can also be alternated, changed, rearranged, decomposed, combined, or deleted.

[0114] The above are only embodiments of the present invention, and thus do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A VR video processing method, characterized in that, the method comprises the following steps: Step S1, a first user collects spatial scene information of a first space to generate a first VR video, and the first VR video includes first identification information and first acquisition time information of the first space; Step S2, a second user collects video information or image information of a second space, and the video information or the image information includes second identification information and second acquisition time information of the second space; Step S3, determine whether the second identification information is the same as the first identification information. If they are the same, execute Step S4; otherwise, save the video information or the image information; Step S4, determine whether the second acquisition time information is later than the first acquisition time information. If it is, execute Step S5; otherwise, end the process; Step S5, use the video information or the image information to update the first VR video, so that when a third user stays at a certain perspective scene during the process of watching the first VR video, a scene transition from the scene at the first acquisition time to the scene at the second acquisition time is shown.

2. The VR video processing method according to claim 1, characterized in that, after determining whether the second identification information is the same as the first identification information and if they are the same, it further includes: establish a timeline according to the first acquisition time information and the second acquisition time information; establish an association relationship between the first VR video and the timeline; establish an association relationship between the video information or the image information and the timeline.

3. The VR video processing method according to claim 2, characterized in that, using the video information to update the first VR video further includes: obtain multiple key perspective information in the first VR video, and obtain corresponding perspective scenes in the first VR video according to the multiple key perspective information; search for perspective information matching the key perspective information from the video information, and establish a corresponding relationship between each key perspective information and the matching perspective information; update the corresponding perspective scenes in the first VR video with the perspective scenes associated with the matching perspective information according to the corresponding relationship.

4. The VR video processing method according to claim 3, characterized in that, searching for perspective information matching the key perspective information from the video information and establishing a corresponding relationship between each key perspective information and the matching perspective information further includes: mark the key perspective information at the position of the first acquisition time on the timeline to obtain multiple key perspective marks; search for perspective information matching the key perspective information from the video information, and mark the matching perspective information at the position of the second acquisition time on the timeline to obtain multiple matching perspective marks; establish a corresponding relationship between each matching perspective mark and the corresponding key perspective mark.

5. The VR video processing method according to claim 2, characterized in that, using the image information to update the first VR video further includes: Obtain multiple key video frames in the first VR video, and search for image information matching the key video frames from the image information; Update the corresponding key video frames in the first VR video with the matching images.

6. The VR video processing method according to claim 5, wherein, Updating the corresponding key video frames in the first VR video with the matching images further includes: Mark the key video frames at the first acquisition time position of the timeline to obtain multiple key video frame identifiers; Search for image information matching the key video frame information from the image information, and mark the matching image information at the second acquisition time position of the timeline to obtain multiple matching image identifiers; Establish a corresponding relationship between each of the matching image identifiers and the corresponding key video frame identifiers.

7. A VR video processing device, wherein, The device includes the following modules: A VR video acquisition module, configured to acquire spatial scene information of a first space for a first user to generate a first VR video, where the first VR video includes first identification information and first acquisition time information of the first space; An update data acquisition module, configured to acquire video information or image information of a second space for a second user, where the video information or image information includes second identification information and second acquisition time information of the second space; A first judgment module, configured to judge whether the second identification information is the same as the first identification information. If they are the same, execute the second judgment module; otherwise, save the video information or the image information; A second judgment module, configured to judge whether the second acquisition time information is later than the first acquisition time information. If so, execute the VR video update module; A VR video update module, which updates the first VR video with the video information or the image information, so that when a third user stays at a certain perspective scene during the process of watching the first VR video, a scene transition from the scene at the first acquisition time to the scene at the second acquisition time is shown.

8. The VR video processing device according to claim 7, wherein, The first judgment module further includes: Establish a timeline according to the first acquisition time information and the second acquisition time information; Establish an association relationship between the first VR video and the timeline; Establish an association relationship between the video information or the image information and the timeline.

9. A VR video processing system, wherein, The system executes the VR video processing method according to any one of claims 1-7.

10. A computer-readable storage medium, the computer-readable storage medium stores a computer program, wherein, The computer program executes the VR video processing method according to any one of claims 1-7.

Citation Information

Patent Citations

  • VR house viewing method and device, computer equipment and storage medium

    CN108492379A

  • Display method and device, terminal equipment and medium

    CN110930220A