Video generation method and apparatus, electronic device, and storage medium

By generating and combining videos with apartment feature tags, the problem of missing features in apartment displays was solved, enabling users to fully understand the features of the apartments.

CN115714840BActive Publication Date: 2026-04-28BEIJING CHENGSHI WANGLIN INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CHENGSHI WANGLIN INFORMATION TECH CO LTD
Filing Date
2022-10-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The existing method of displaying apartment layouts requires users to actively identify the features of the apartment, which can easily lead to omissions of features and make it impossible to accurately obtain the features of the apartment.

Method used

By determining N feature labels based on the feature parameters of the target floor plan in the target dimension, a first video and a second video corresponding to the N feature labels are generated, and they are combined into a guided video to showcase the features of the floor plan.

Benefits of technology

This allows users to efficiently and accurately understand the features of a unit without having to actively identify them, thus improving the information acquisition experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115714840B_ABST
    Figure CN115714840B_ABST
Patent Text Reader

Abstract

The application provides a video generation method and device, electronic equipment and storage medium, the method comprises the following steps: determining N feature labels corresponding to a target house type drawing according to feature parameters of the target house type drawing in a target dimension, N is an integer greater than or equal to 1; generating a first video according to the N feature labels, the target house type drawing, a community panoramic dynamic drawing and a first explanation audio; for each feature label, generating a second video corresponding to the feature label according to a video frame sequence corresponding to the feature label and a second explanation audio; combining the first video and the N second videos to generate a tour video corresponding to the target house type drawing. The application can generate a tour video based on the house type characteristics of the target house type drawing, facilitate the user to efficiently and accurately understand the target house type drawing as a whole and understand different features of the target house type drawing based on each feature label, so that the user can relatively comprehensively understand the house type characteristics, and the information acquisition experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a video generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] During the property transaction process, users need to understand the property information. Online property viewing is favored by many users due to its convenience and efficiency. Therefore, online property viewing plays an increasingly important role in the property viewing scenario.

[0003] As a key focus for users in property transactions, the floor plan of a property is usually displayed in online property viewing scenarios so that users can view the floor plan information and extract its key features.

[0004] However, the existing method of displaying apartment layouts requires users to rely on their experience to identify the possible features of the layout, which can easily lead to the omission of features. In addition, the criteria for determining the features of some layouts are complex, making it impossible for users to obtain accurate features.

[0005] It is evident that the existing method of displaying apartment layouts requires users to actively identify apartment features, which can easily lead to omissions of features and an inability to accurately obtain apartment features. Summary of the Invention

[0006] This application provides a video generation method, apparatus, electronic device, and storage medium to solve the problems of existing apartment layout display methods that require users to actively identify apartment layout features, which can easily lead to feature omissions and inaccurate acquisition of apartment layout features.

[0007] In a first aspect, embodiments of this application provide a video generation method, including:

[0008] Based on the feature parameters of the target floor plan in the target dimension, determine N feature labels corresponding to the target floor plan, where N is an integer greater than or equal to 1;

[0009] A first video is generated based on the N feature tags, the target floor plan, the panoramic dynamic image of the community, and the first explanatory audio.

[0010] For each feature tag, a second video corresponding to the feature tag is generated based on the video frame sequence corresponding to the feature tag and the second narration audio.

[0011] The first video is combined with N second videos to generate a guided tour video corresponding to the target floor plan.

[0012] Secondly, embodiments of this application provide a video generation apparatus, including:

[0013] The determination module is used to determine N feature labels corresponding to the target floor plan based on the feature parameters of the target floor plan in the target dimension, where N is an integer greater than or equal to 1;

[0014] The first generation module is used to generate a first video based on the N feature tags, the target floor plan, the panoramic dynamic map of the community, and the first explanatory audio.

[0015] The second generation module is used to generate a second video corresponding to each feature label based on the video frame sequence corresponding to the feature label and the second narration audio.

[0016] The third generation module is used to combine the first video with N second videos to generate a tour video corresponding to the target floor plan.

[0017] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the video generation method as described in the first aspect above.

[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the video generation method described in the first aspect above.

[0019] The technical solution of this application embodiment determines N feature tags corresponding to the target floor plan based on the feature parameters of the target floor plan in the target dimension. A first video is generated to provide an overall introduction to the target floor plan based on the N feature tags, the target floor plan, a panoramic dynamic image of the community to which the target floor plan belongs, and a first explanatory audio. For each of the N feature tags, a second video is generated to individually introduce the floor plan features of the feature tag based on the corresponding video frame sequence and the second explanatory audio. The first video and the N second videos are then stitched together to generate a guide video for the target floor plan. This allows for the generation of guide videos based on the floor plan features of the target floor plan, enabling users to efficiently and accurately gain an overall understanding of the target floor plan and to understand the different features of the target floor plan based on each feature tag. This provides users with a relatively comprehensive understanding of the floor plan features and improves their information acquisition experience. Attached Figure Description

[0020] Figure 1 A schematic diagram illustrating the video generation method provided in an embodiment of this application;

[0021] Figure 2This is a schematic diagram illustrating video frame images including a panoramic view of the community and a target apartment floor plan provided in an embodiment of this application.

[0022] Figure 3 This is a schematic diagram illustrating a video frame image in a spatially characteristic video frame sequence provided in an embodiment of this application;

[0023] Figure 4 This is a schematic diagram illustrating how the label provided in this application displays a video frame image in a video frame sequence;

[0024] Figure 5a This is a schematic diagram illustrating the addition of subtitle information to a contour marker diagram according to an embodiment of this application;

[0025] Figure 5b This is a schematic diagram illustrating the addition of subtitle information and target floor plan to a real-world image, as provided in the embodiments of this application.

[0026] Figure 6a and Figure 6b This is a schematic diagram illustrating the marking of the entryway in the target floor plan according to an embodiment of this application;

[0027] Figure 6c and Figure 6d This is a schematic diagram illustrating the marking of a living room with a balcony in a target floor plan, as provided in an embodiment of this application.

[0028] Figure 7 This is a schematic diagram illustrating the target floor plan for adding circulation routes according to an embodiment of this application;

[0029] Figure 8a This is a schematic diagram illustrating the addition of subtitle information to a two-dimensional ventilation effect diagram provided in an embodiment of this application.

[0030] Figure 8b This is a schematic diagram illustrating the addition of subtitle information to a three-dimensional ventilation effect diagram provided in an embodiment of this application.

[0031] Figure 9 A schematic diagram illustrating the window marker image for adding caption information provided in an embodiment of this application;

[0032] Figure 10 A schematic diagram illustrating the two-dimensional lighting effect of adding subtitle information provided in the embodiments of this application;

[0033] Figure 11 This is a schematic diagram illustrating the video generation apparatus provided in an embodiment of this application;

[0034] Figure 12 This is a schematic diagram of the electronic device structure provided in the embodiments of this application. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0036] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0037] In the various embodiments of this application, it should be understood that the sequence number of each process described below does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0038] To address the problems of existing floor plan display methods requiring users to actively identify floor plan features, which can easily lead to omissions and inaccurate acquisition of features, this application provides a video generation method. By analyzing the floor plan and extracting relevant floor plan tags, the method can obtain the features of the floor plan. Based on the features and explanatory audio, a guide video corresponding to the target floor plan is generated and displayed. This allows users to efficiently and accurately understand the features of the floor plan without having to identify the floor plan themselves, thus improving the user's information acquisition experience.

[0039] The video generation method provided in this application is described below; see [link to relevant documentation]. Figure 1 As shown, the method includes the following steps:

[0040] Step 101: Based on the feature parameters of the target floor plan in the target dimension, determine the N feature labels corresponding to the target floor plan, where N is an integer greater than or equal to 1.

[0041] The video generation method of this application first obtains the feature parameters corresponding to the target floor plan in the target dimension. The target floor plan can be any one of multiple floor plans. After obtaining the feature parameters of the target floor plan in the target dimension, N feature labels corresponding to the target floor plan in the target dimension are determined based on the obtained feature parameters. By analyzing the target floor plan to extract feature parameters, and generating corresponding feature labels based on the extracted feature parameters, the unique features of the target floor plan are obtained.

[0042] When determining the feature label corresponding to the target floor plan based on the feature parameters of the target floor plan in the target dimension, one or more feature labels corresponding to the target floor plan in the target dimension can be determined based on one or more feature parameters, so as to obtain the floor plan features corresponding to the target floor plan based on at least one feature label.

[0043] In this application embodiment, the target floor plan can be a two-dimensional or three-dimensional floor plan displayed on a 3D digital interface. The 3D digital interface can be an interactive interface such as Virtual Reality (VR), Augmented Reality (AR), or panoramic. When the target floor plan consists of both two-dimensional and three-dimensional floor plans of the same apartment type, feature parameters of the target floor plan can be obtained based on both the two-dimensional and three-dimensional floor plans to acquire relatively accurate and comprehensive feature parameters, and then feature labels can be determined based on these feature parameters.

[0044] Step 102: Generate a first video based on the N feature tags, the target floor plan, the panoramic dynamic map of the community, and the first explanatory audio.

[0045] After obtaining N feature tags corresponding to the target floor plan, a first video can be generated based on the N feature tags, the target floor plan, the panoramic dynamic image of the community to which the target floor plan belongs, and the first explanatory audio. The first explanatory audio is used to provide an overall introduction to the target floor plan, which may include an introduction to the features of the floor plan and other relevant information. The first explanatory audio is matched with the visuals corresponding to the N feature tags, the target floor plan, and the panoramic dynamic image of the community to generate a first video that combines visuals and audio.

[0046] Among them, N feature labels can be added to the target floor plan to work with it, or they can be added to the community panoramic dynamic image to work with it.

[0047] Step 103: For each feature tag, generate a second video corresponding to the feature tag based on the video frame sequence corresponding to the feature tag and the second narration audio.

[0048] For each of the N feature labels corresponding to the target floor plan, a second video corresponding to the feature label is generated based on the video frame sequence corresponding to the feature label and the second explanatory audio. The video frame sequence corresponding to the feature label is used to introduce the floor plan features corresponding to the feature label, and the second explanatory audio corresponding to the feature label is matched with the video frame sequence corresponding to the feature label. By matching the video frame sequence corresponding to the feature label with the second explanatory audio corresponding to the feature label, the features of the floor plan corresponding to the feature label are introduced in both visual and audio dimensions.

[0049] By generating a corresponding second video for each feature label, the apartment layout features corresponding to different feature labels can be introduced separately. Furthermore, when N is greater than or equal to 2, the N feature labels can be prioritized, and second videos corresponding to each feature label can be generated sequentially in descending order of priority.

[0050] Step 104: Combine the first video with N second videos to generate a guided tour video corresponding to the target floor plan.

[0051] After generating a first video providing an overall introduction to the target floor plan and N second videos corresponding to N feature tags, the first video and the N second videos are combined to generate a guide video that provides an overall introduction to the target floor plan and a separate introduction to each feature tag. The process of combining the first video and the N second videos can be understood as splicing them together. During splicing, the first video can be spliced ​​first, followed by the N second videos, generating a guide video that first shows the overall situation of the target floor plan and then the floor plan features of each feature tag; alternatively, the N second videos can be spliced ​​first, followed by the first video, generating a guide video that first shows the floor plan features of each feature tag and then the overall situation of the target floor plan; or, when N is greater than or equal to 2, the first video can be inserted into the N second videos for splicing. Other splicing methods are also possible, but will not be elaborated upon here.

[0052] The above-described implementation process of this application determines N feature labels corresponding to the target floor plan based on the feature parameters of the target floor plan in the target dimension. A first video is generated to provide an overall introduction to the target floor plan based on the N feature labels, the target floor plan, a panoramic dynamic image of the community to which the target floor plan belongs, and a first explanatory audio. For each of the N feature labels, a second video is generated to individually introduce the floor plan features of that feature label based on the corresponding video frame sequence and the second explanatory audio. The first video and the N second videos are then stitched together to generate a guided tour video of the target floor plan. This allows for the generation of guided tour videos based on the floor plan features of the target floor plan, enabling users to efficiently and accurately gain an overall understanding of the target floor plan and to understand the different features of the target floor plan based on each feature label. This provides users with a relatively comprehensive understanding of the floor plan features and improves their information acquisition experience.

[0053] In one embodiment of this application, the target dimension includes at least one of spatial scale dimension and light transmission dimension; the N feature labels include at least one of the following associated with the spatial scale dimension: high space utilization feature label, square apartment layout feature label, structural feature label, active and quiet zone feature label, reasonable circulation feature label, and reasonable width and depth feature label, and / or, the N feature labels include at least one of the following associated with the light transmission dimension: ventilation feature label, apartment type feature label, and reasonable lighting feature label.

[0054] In this embodiment, the target dimension includes at least one of the spatial scale dimension and the transparency and lighting dimension. The feature parameters of the target floor plan in the spatial scale dimension are associated with the features of the floor plan in terms of floor plan structure. The feature parameters of the target floor plan in the transparency and lighting dimension are associated with the features of the floor plan in terms of ventilation, lighting and orientation.

[0055] When the target dimension includes spatial scale, one or more feature parameters corresponding to the target floor plan in this dimension can be obtained. These feature parameters may include space utilization, floor plan squareness, room structure, distribution of active and quiet areas, circulation layout, and width and depth, among others. When the target dimension includes ventilation and lighting, one or more feature parameters corresponding to these dimensions can be obtained. These feature parameters may include ventilation conditions, floor plan type, and lighting conditions, among others. When the target dimensions include both spatial scale and ventilation / lighting dimensions, one or more feature parameters corresponding to each dimension can be obtained.

[0056] Accordingly, based on the characteristic parameters of the target floor plan in the target dimension, the N characteristic labels determined may include at least one of the following: high space utilization characteristic label associated with the spatial scale dimension, square floor plan characteristic label, structural characteristic label, active and quiet zone characteristic label, reasonable circulation characteristic label, and reasonable width and depth characteristic label, and / or at least one of the following: ventilation characteristic label associated with the transparency and lighting dimension, floor plan category characteristic label, and reasonable lighting characteristic label.

[0057] When generating feature labels based on the feature parameters of the spatial scale dimension, a high space utilization feature label can be generated when the space utilization rate of the target floor plan meets the first preset condition; a square floor plan feature label can be generated when the squareness of the target floor plan meets the second preset condition; a structural feature label can be generated when certain rooms in the target floor plan meet the corresponding preset structural and functional features; a dynamic and static zoning feature label can be generated when the room distribution in the target floor plan meets the dynamic and static zoning; a reasonable circulation feature label can be generated when the room distribution in the target floor plan meets the reasonable circulation layout; and a reasonable width and depth feature label can be generated when the number of rooms in the target floor plan that meet the corresponding width and depth conditions is greater than a preset threshold.

[0058] When generating feature labels based on the characteristic parameters of the transparency and lighting dimension, at least one of the following can be generated: ventilation feature label, apartment type feature label, and lighting feature label. Specifically, a ventilation feature label can be generated based on ventilation conditions; an apartment type feature label can be generated based on apartment type; and a lighting feature label can be generated based on lighting conditions.

[0059] When generating feature tags based on ventilation conditions, if the ventilation condition indicates a direct ventilation type, a direct ventilation feature tag is generated; if the ventilation condition indicates a two-sided ventilation type, a two-sided ventilation feature tag is generated; and if the ventilation condition indicates a one-sided ventilation type, a one-sided ventilation feature tag is generated. When generating unit type feature tags based on unit type, one of the following feature tags can be generated based on the unit type corresponding to the target unit type: fully illuminated unit, north-south facing unit, south-facing unit, southeast-facing unit, southwest-facing unit, east-facing unit, west-facing unit, east-west facing unit, northeast-facing unit, northwest-facing unit, and north-facing unit. When generating lighting feature tags based on daylighting conditions, a reasonable daylighting feature tag can be generated if the target unit type meets reasonable daylighting conditions.

[0060] The above-described implementation process of this application can obtain the feature parameters of the target floor plan in the spatial scale dimension and / or the transparency and lighting dimension, generate feature labels corresponding to the spatial scale dimension and / or the transparency and lighting dimension based on the obtained feature parameters, and generate a guide video based on the feature labels corresponding to the spatial scale dimension and / or the transparency and lighting dimension, so as to facilitate users to understand the target floor plan in the spatial scale dimension and / or the transparency and lighting dimension.

[0061] The process of generating the first video is described below. When generating the first video based on the N feature tags, the target floor plan, the panoramic dynamic image of the community, and the first narration audio, it includes:

[0062] Based on the N feature labels, the target floor plan, the community panoramic dynamic image, and the first text corpus, a sequence of overall introduction video frames corresponding to the target floor plan is generated, and the community panoramic dynamic image is the panoramic dynamic image corresponding to the target community to which the target floor plan belongs.

[0063] Based on the target floor plan, the panoramic dynamic image of the community, and the second text corpus, a spatially distinctive video frame sequence is generated;

[0064] Based on the target floor plan and the N feature labels, generate a sequence of labeled display video frames;

[0065] Based on the target floor plan, motion effect rules, and third-party text corpus, generate a sequence of video frames for displaying the floor plan.

[0066] The first video frame sequence is generated by combining the overall introduction video frame sequence, the spatial feature video frame sequence, the tag display video frame sequence, and the floor plan display video frame sequence.

[0067] The first video frame sequence and the first narration audio are matched to generate the first video. The text corpus corresponding to the first narration audio includes the first text corpus, the second text corpus, the tag text corpus that matches the N feature tags, and the third text corpus.

[0068] When generating the first video, a sequence of overall introduction video frames corresponding to the target floor plan can be generated based on N feature tags, the target floor plan, a panoramic dynamic image of the community, and the first text corpus. The first text corpus can be understood as the subtitles in the overall introduction video frame sequence, and the overall introduction screen about the target floor plan formed by the combination of the target floor plan, feature tags, and the panoramic dynamic image of the community is adapted to the first text corpus. In this embodiment, the panoramic dynamic image of the community is a panoramic dynamic image of the target community to which the property corresponding to the target floor plan belongs.

[0069] When generating the first video, a spatially distinctive video frame sequence needs to be generated based on the target floor plan, the panoramic dynamic image of the community, and the second text corpus. Correspondingly, the second text corpus can be understood as the subtitles in the spatially distinctive video frame sequence, and the spatially distinctive visuals formed by the combination of the target floor plan and the panoramic dynamic image of the community are adapted to the second text corpus.

[0070] When generating the first video, it is also necessary to generate a sequence of video frames displaying tags and a sequence of video frames displaying the floor plan. When generating the sequence of video frames displaying tags, it is based on the target floor plan and N feature tags. This sequence of video frames can include at least one video frame image. When generating the sequence of video frames displaying the floor plan, it is based on the target floor plan, the corresponding motion effect rules, and a third text corpus. The third text corpus can be understood as the subtitles in the sequence of video frames displaying the floor plan, and the visuals formed by the target floor plan and the motion effect rules are adapted to the third text corpus.

[0071] After generating the overall introduction video frame sequence, the spatial feature video frame sequence, the tag display video frame sequence, and the floor plan display video frame sequence, these sequences can be combined to generate the first video frame sequence. The process of combining multiple video frame sequences can be understood as splicing video frame sequences. During splicing, the sequences can be spliced ​​according to a pre-set splicing order or according to the generation order of the video frame sequences.

[0072] After generating the first video frame sequence by splicing, the first video frame sequence is matched with the first narration audio to generate the first video, which includes both visuals and audio. Matching the first video frame sequence with the first narration audio means merging the first video frame sequence with the first narration audio, achieving coordination between visuals and audio. It can also be understood as adding the first narration audio to the first video frame sequence to generate the first video.

[0073] The text corpus corresponding to the first audio explanation includes a first text corpus, a second text corpus, a tagged text corpus matched with N feature labels, and a third text corpus. The first text corpus, the second text corpus, the tagged text corpus, and the third text corpus can be used as subtitles in the video frame sequence. By matching the first audio explanation with the first video frame sequence, the audio explanation can be adapted to the text corpus, ensuring that the audio and subtitles match when the audio explanation is played.

[0074] It should be noted that the audio narration order corresponding to the first audio explanation is a pre-set order. Therefore, when splicing the video frame sequence, it is necessary to splice the overall introduction video frame sequence, spatial feature video frame sequence, label display video frame sequence, and floor plan display video frame sequence according to the pre-set splicing order that matches the audio narration order, so as to ensure that the display of the text corpus is compatible with the playback of the audio explanation.

[0075] In this embodiment, when splicing video frame sequences, they can be spliced ​​in the following order: general introduction video frame sequence, spatial feature video frame sequence, tag display video frame sequence, and floor plan display video frame sequence. Correspondingly, the text corpus corresponding to the first explanation audio can be arranged in the following order: first text corpus, second text corpus, tag text corpus, and third text corpus.

[0076] The above-described implementation process of this application generates a sequence of video frames for overall introduction, spatial features, tags, and floor plan display based on N feature tags, the target floor plan, a panoramic dynamic image of the community, and text corpus. The generated video frame sequences are then spliced ​​together according to a preset splicing order to generate a first video frame sequence. This first video frame sequence is then matched with a first audio explanation (the explanation order matches the splicing order) to generate a first video. Based on the splicing of the video frame sequences and the audio-visual matching, a first video for providing an overall introduction to the target floor plan can be generated, allowing users to gain a comprehensive understanding of the target floor plan during subsequent presentations.

[0077] The process of generating the overall introduction video frame sequence, the spatial feature video frame sequence, the tag display video frame sequence, and the floor plan display video frame sequence is described below.

[0078] For generating a sequence of overall introductory video frames, when generating the sequence of overall introductory video frames corresponding to the target floor plan based on the N feature labels, the target floor plan, the panoramic dynamic image of the community, and the first text corpus, one of the following schemes may be included:

[0079] The N feature labels and the first text corpus are added to the panoramic dynamic map of the community to generate the first target dynamic map. Each frame of the first target dynamic map is stitched together with the target floor plan to generate the overall introduction video frame sequence.

[0080] Each frame of the panoramic dynamic image of the community is stitched together with the target floor plan to generate a second target dynamic image. The N feature labels and the first text corpus are added to the panoramic area of ​​the community in the second target dynamic image to generate the overall introduction video frame sequence.

[0081] The duration of the overall introduction video frame sequence is equal to the duration of the first text corpus.

[0082] When generating the overall introduction video frame sequence, N feature labels and the first text corpus can be added to the cell panoramic dynamic image to generate the first target dynamic image. Since the cell panoramic dynamic image includes multiple static images, N feature labels and at least part of the content from the first text corpus are added to each frame. The N feature labels can be displayed in a preset style, such as displaying the N feature labels at a first size and 30% transparency. The number of static image frames corresponding to the first target dynamic image is equal to the number of static image frames corresponding to the cell panoramic dynamic image, and each static image frame in the first target dynamic image corresponds to N feature labels. The subtitles corresponding to each static image frame are at least part of the content from the first text corpus. The cell panoramic dynamic image serves as the background for the subtitles and the N feature labels.

[0083] After generating the first target animated image, it is stitched together with the target floor plan to generate a sequence of overall introduction video frames. Since the first target animated image corresponds to multiple static images, the target floor plan is stitched together with each static image frame during the stitching process. Then, the overall introduction video frame sequence is generated based on the stitched image. The number of frames in the overall introduction video frame sequence is the same as the number of frames in the panoramic animated image of the community.

[0084] When generating the overall introduction video frame sequence, each static image corresponding to the panoramic dynamic image of the community can be stitched together with the target floor plan to generate a second target dynamic image. The number of static image frames corresponding to the second target dynamic image is equal to the number of static image frames corresponding to the panoramic dynamic image of the community. Each static image frame in the second target dynamic image includes both the target floor plan and the panoramic image of the community. After generating the second target dynamic image, N feature labels and a first text corpus can be added to the panoramic area of ​​the community in the second target dynamic image to generate the overall introduction video frame sequence. When adding N feature labels and the first text corpus to the panoramic area of ​​the community in the second target dynamic image, for each static image in the second target dynamic image, N feature labels and at least a portion of the first text corpus are added, with the added first text corpus serving as subtitle information.

[0085] The duration of the overall introduction video frame sequence is equal to the duration of the first text corpus. The duration of the first text corpus is the broadcast duration of the first text corpus, which can also be understood as the total display duration of the first text corpus. The first text corpus serves as subtitles in the overall introduction video frame sequence. The display duration of each line of subtitles is equal to the audio broadcast duration of that line. The first text corpus can correspond to one or more lines of subtitles. In the case of multiple lines of subtitles, each line of subtitles corresponds to one or more static images. The static images corresponding to different lines of subtitles can be different. Each line of subtitles is a portion of the content in the first text corpus.

[0086] When stitching together the target floor plan, you can add relevant floor plan descriptions, such as the layout and corresponding area information. See also... Figure 2 The image shown is a schematic diagram of a video frame in the overall introduction video frame sequence. The video frame image includes a target floor plan with relevant descriptions (98.5 square meters, 3 bedrooms, 2 living rooms, 2 bathrooms) and a panoramic static image of the community with feature tags (high utilization rate, independent entrance hall, living room with balcony, bright floor plan, direct ventilation, good lighting) and subtitle information (floor plan analysis of 3 bedrooms, 2 living rooms, 2 bathrooms).

[0087] The above-described implementation process of this application involves using a panoramic dynamic image of the community as a background, adding feature labels and a first text corpus to the panoramic dynamic image of the community to generate a first target dynamic image, and then stitching each frame of the first target dynamic image with the target floor plan to generate a sequence of overall introduction video frames. This can be done by adding the features first and then stitching the frames together. Similarly, by stitching each frame of the panoramic dynamic image of the community with the target floor plan to generate a second target dynamic image, and then adding N feature labels and the first text corpus to the second target dynamic image to generate a sequence of overall introduction video frames, this can also be done by stitching the features first and then adding the features first.

[0088] By generating a video frame sequence that introduces the overall situation of the target floor plan using any of the above methods, users can gain a relatively comprehensive understanding of the target floor plan, the features of the corresponding floor plan, and the situation of the community to which the target floor plan belongs.

[0089] For generating spatially distinctive video frame sequences, the process of generating such sequences based on the target floor plan, the panoramic dynamic image of the community, and the second text corpus includes:

[0090] Using the panoramic dynamic image of the community as the background and the target apartment floor plan as the main body, a sequence of spatially characteristic video frames with the second text corpus added is generated.

[0091] The duration of the spatially characteristic video frame sequence is equal to the duration of the second text corpus.

[0092] When generating a spatially distinctive video frame sequence, a panoramic dynamic image of the community can be used as the background, and the target floor plan can be used as the main body. The target floor plan is added to the panoramic dynamic image of the community to generate an image sequence. Then, the second text corpus is added to the generated image sequence to generate a spatially distinctive video frame sequence.

[0093] When adding the second text corpus, it can be added to the background area; that is, the panoramic dynamic image of the community can be used as the background for the second text corpus. It should be noted that, in order to show that the panoramic dynamic image of the community is the background, a transparent overlay can be superimposed on the panoramic dynamic image, such as a black overlay with 50% transparency, and then the target apartment floor plan and the second text corpus can be added on the overlay.

[0094] The duration of the spatially-featured video frame sequence is equal to the duration of the second text corpus. The duration of the second text corpus is the broadcast duration of the second text corpus, which can also be understood as the total display duration of the second text corpus. The second text corpus serves as subtitles in the spatially-featured video frame sequence, and the display duration of each line of subtitles is equal to the audio broadcast duration of that line of subtitles. The second text corpus can correspond to one or more lines of subtitles.

[0095] See Figure 3 The image shown is a schematic diagram of a video frame in a spatial feature video frame sequence. In the video frame image, a panoramic view of the community serves as the background, and a transparent overlay is set on the background. The target apartment floor plan and subtitle information (Let's look at the spatial features of this apartment) are added on the transparent overlay. The subtitle information is at least part of the content of the second text corpus.

[0096] The above-described implementation process of this application, by using a panoramic dynamic map of the community as a background and adding the target floor plan and second text corpus to the panoramic dynamic map of the community, allows users to understand the spatial structure of the target floor plan while also understanding the community situation of the community to which the target floor plan belongs through the background.

[0097] For generating a video frame sequence for displaying tags, the process of generating the video frame sequence for displaying tags based on the target floor plan and the N feature tags includes:

[0098] The N feature labels are sequentially added to the target floor plan to generate the label display video frame sequence;

[0099] Wherein, the duration of the video frame sequence displayed by the tag is equal to the duration of the tag text corpus.

[0100] When generating the video frame sequence for label display, N pre-generated feature labels are sequentially added to the target floor plan. Multiple images are generated by adding feature labels one by one, thus creating the video frame sequence for label display. For cases where N is greater than 1, the display styles corresponding to the N feature labels can be the same, or different feature labels can have different display styles. When adding feature labels to the target floor plan sequentially, they can be added based on a preset label priority; that is, higher priority labels can be added earlier.

[0101] The duration of the video frame sequence displayed by the tags is equal to the duration of the tag text corpus associated with the N feature tags. The duration of the tag text corpus associated with the N feature tags is the broadcast duration of the tag text corpus, which can also be understood as the total display duration of the N feature tags displayed in sequence.

[0102] It should be noted that the background area corresponding to the target floor plan can be a non-transparent overlay. N feature labels can correspond to one add box; that is, the feature labels are displayed within the add box, which is superimposed on the target floor plan. The add box can have a preset transparency to avoid obscuring the target floor plan. The size of the add box can be adjusted, and correspondingly, the size of the feature labels can also be adjusted. When there are too many feature labels, the size of the add box can be adjusted automatically or manually, as can the size of the feature labels.

[0103] See Figure 4 The image shown is a schematic diagram of a video frame image in a video frame sequence for label display. In the video frame image, a non-transparent overlay serves as the background, and the target floor plan is displayed on the overlay. The addition boxes corresponding to N feature labels are superimposed on the target floor plan with preset transparency, and the N feature labels (high utilization rate, independent entrance hall, living room with balcony, reasonable width and depth) are displayed in the addition boxes.

[0104] The above implementation process of this application, by adding N feature tags sequentially to the target floor plan, can attract users' attention to the feature tags corresponding to the target floor plan by displaying the feature tags in sequence, making it easier for users to understand the features of the floor plan.

[0105] For generating a sequence of video frames for displaying a floor plan, the process of generating the sequence based on the target floor plan, motion effect rules, and third-party text corpus includes:

[0106] Generate an animated floor plan based on the animation rules and the target floor plan;

[0107] The third text corpus is added to the animated floor plan to generate a sequence of video frames displaying the floor plan.

[0108] The duration of the video frame sequence displaying the floor plan is equal to the duration of the third text corpus.

[0109] When generating a video frame sequence for displaying floor plans, the target floor plan is processed according to motion effect rules to generate an animated floor plan. These rules can involve scaling each room in the target floor plan, with different rooms having the same or different scaling ratios. Different rooms can scale synchronously or sequentially according to a preset order. The generated animated floor plan can include multiple target floor plans in different scaling states. Third-party text corpus is added to the generated animated floor plan to produce the video frame sequence for displaying the floor plan.

[0110] By using animation rules to scale the target floor plan and generate an animated floor plan, users can be drawn to the rooms in the target floor plan by changing the display style of each room.

[0111] The duration of the video frame sequence displayed in the floor plan is equal to the duration of the third text corpus. The duration of the third text corpus is the broadcast duration of the third text corpus, which can also be understood as the total display duration of the third text corpus.

[0112] The above implementation process of this application generates a floor plan animation based on motion effect rules and the target floor plan, adds third-party text corpus to the floor plan animation, and generates a sequence of floor plan display video frames. This can attract users' attention to the target floor plan by changing the display style of each room in the target floor plan.

[0113] In one embodiment of this application, when the value of N is greater than or equal to 2, the first video is combined with N second videos to generate a guide video corresponding to the target floor plan, including: sorting the N second videos according to the priority of feature tags; and splicing the first video and the sorted N second videos to generate the guide video.

[0114] When there are two or more generated feature labels, when concatenating the first video with N second videos, the N second videos can be sorted according to the priority of the N feature labels, with higher priority videos appearing first. After sorting the N second videos, the first video is concatenated with the sorted N second videos. The concatenation can be performed in the following order: first video at the beginning, followed by the N sorted second videos in sequence; first video at the end, followed by the N sorted second videos in sequence; or the first video can be inserted into the N sorted second videos.

[0115] By sorting N second videos and generating a guide video based on the first video and the sorted N second videos, the guide video can be displayed according to the priority of the feature tags when showing the guide video.

[0116] The following describes the process of generating the second video corresponding to the feature label. The process of generating the second video corresponding to the feature label based on the video frame sequence corresponding to the feature label and the second narration audio includes:

[0117] The text corpus corresponding to the feature label is added to the target image sequence corresponding to the feature label to generate the video frame sequence corresponding to the feature label;

[0118] The video frame sequence corresponding to the feature label and the second audio explanation are matched to generate the second video corresponding to the feature label.

[0119] The second audio explanation corresponding to the feature label is adapted to the text corpus corresponding to the feature label.

[0120] When generating the second video corresponding to the feature label, the text corpus corresponding to the feature label can be added to the target image sequence corresponding to the feature label to generate the video frame sequence corresponding to the feature label. Then, the video frame sequence corresponding to the feature label is matched with the second narration audio corresponding to the feature label to generate the second video corresponding to the feature label.

[0121] The second audio explanation corresponding to the feature tag is matched with the text corpus corresponding to the feature tag. Matching here means that the audio content of the second audio explanation is identical to the text content of the text corpus. The duration of the second video corresponding to the feature tag is equal to the total display duration of the text corpus corresponding to the feature tag, which is the playback duration of the second audio explanation. The text corpus corresponding to the feature tag consists of the subtitles in the second video corresponding to the feature tag, and the display duration of each line of subtitles is equal to the audio playback duration of that line of subtitles.

[0122] The following section introduces the relationship between feature labels and spatial scale dimensions. When feature labels are associated with spatial scale dimensions, feature labels can include at least one of the following: high space utilization feature labels, square apartment layout feature labels, structural feature labels, active and quiet zone feature labels, reasonable circulation feature labels, and reasonable width and depth feature labels.

[0123] Optionally, when the feature label is the high space utilization feature label, the target image sequence includes an outline marker animation and / or a real-scene animation corresponding to the target floor plan. The outline marker animation corresponding to the target floor plan may include multiple outline marker images corresponding to the target floor plan, and each outline marker image has a different outline marker state. The marker style of the outline marker image can be a highlight style, a specific color, etc. The real-scene animation corresponding to the target floor plan is generated by taking real-scene photos of the property corresponding to the target floor plan.

[0124] When generating the second video corresponding to the high space utilization feature label, the space utilization text corpus corresponding to the high space utilization feature label is added to the outline marker animation and / or real-scene animation corresponding to the target floor plan to generate the video frame sequence corresponding to the high space utilization feature label. Then, the generated video frame sequence is matched with the corresponding second narration audio to generate the second video corresponding to the high space utilization feature label.

[0125] Specifically, space utilization text corpus is added as caption information to the outline-marked animated image and / or the real-scene animated image. The background of the outline-marked animated image can be a non-transparent overlay. When adding space utilization text corpus as caption information to the outline-marked animated image corresponding to the target floor plan, a high space utilization feature label can also be added. When adding space utilization text corpus as caption information to the real-scene animated image corresponding to the target floor plan, the target floor plan can also be added to the real-scene animated image. See also Figure 5a The image shown is a specific example of adding caption information and highly utilized feature labels to a contour marker map; see also... Figure 5b The image shown is a specific example of adding caption information and the target floor plan to a real-world image.

[0126] By adding space utilization text corpus to the outline marker animation, users can easily understand the spatial structure of the target floor plan based on the outline marker animation and subtitle information; by adding space utilization text corpus to the real-scene animation, users can easily understand the spatial structure of the target floor plan based on the real-scene information and subtitle information.

[0127] Optionally, if the feature label is the square floor plan feature label, the target image sequence includes the target floor plan. The target image sequence may include one or more frames of images. For the case where the feature label is the square floor plan feature label, the target image sequence may include one frame of the target floor plan or multiple frames of target floor plans in different states.

[0128] When generating the second video corresponding to the square floor plan feature label, the square floor plan text corpus corresponding to the square floor plan feature label is added to the target floor plan, generating a video frame sequence corresponding to the square floor plan feature label. Then, the generated video frame sequence corresponding to the square floor plan feature label is matched with the second narration audio to generate the second video corresponding to the square floor plan feature label. The square floor plan text corpus is added as subtitle information to the target floor plan. When the target image sequence includes multiple frames of target floor plans with different states, square floor plan text corpus is added to each frame of the target floor plan. These different states can be due to differences in display style, such as highlighting the outer contour of one frame of the target floor plan, or using a specific color to mark the room outlines in another frame. Furthermore, the second narration audio corresponding to the square floor plan feature label is adapted to the square floor plan text corpus; that is, the audio content of the second narration audio corresponding to the square floor plan feature label is the same as the text content of the square floor plan text corpus.

[0129] The background of the target floor plan can be a non-transparent overlay. When adding the square floor plan text corpus as subtitle information to the target floor plan, square floor plan feature labels can also be added.

[0130] By adding the text corpus of square floor plans to the target floor plan, it is easier for users to obtain the square floor plan characteristics based on the text corpus when browsing the target floor plan to understand the squareness of the floor plan.

[0131] Optionally, when the feature label is the structural feature label, the target image sequence includes a structural analysis animation and / or a structural real-scene animation corresponding to the structural feature label. The structural analysis animation may include an image sequence corresponding to rooms with specific structures or functions, and each room's image sequence may include multiple frames of analysis images corresponding to that room. These analysis images may be labeled images of the room; the labels may be marking the room's outline, length, width, etc., and a transparent mask may be applied to the room during labeling to distinguish it from other rooms. The structural real-scene animation is generated by taking real-scene photos of rooms with specific structures or functions in the property corresponding to the target floor plan.

[0132] The number of structural feature tags can be one or more. When generating the second video corresponding to the structural feature tags, the structural feature text corpus corresponding to at least one structural feature tag is added to the structural analysis animation and / or structural real-scene animation corresponding to at least one structural feature tag to generate the video frame sequence corresponding to the structural feature tags. Then, the video frame sequence corresponding to the structural feature tags is matched with the second narration audio to generate the second video corresponding to the structural feature tags. The structural feature text corpus is added as subtitle information to the structural analysis animation and / or structural real-scene animation. The second narration audio corresponding to the structural feature tags is adapted to the structural feature text corpus; that is, the audio content of the second narration audio corresponding to the structural feature tags is the same as the text content of the structural feature text corpus.

[0133] The background of the structural analysis animation can be a non-transparent overlay. When adding structural feature text corpus as subtitle information to the structural analysis animation, structural feature tags can also be added. The subtitle information can be added on the overlay or superimposed on the structural analysis animation. When adding structural feature text corpus as subtitle information to the structural reality animation, a target floor plan can also be added. See also Figure 6a and Figure 6b The image shown is a specific example of marking the entryway in the target floor plan; see also Figure 6c as well as Figure 6d The image shows a specific example of marking a living room with a balcony in a target floor plan.

[0134] By adding structural feature text corpus to the structural analysis animation and / or structural scene animation corresponding to the structural feature labels, users can better understand the specific spatial structure of the target floor plan while browsing images, in conjunction with subtitle information.

[0135] Optionally, when the feature label is the dynamic-static partition feature label, the target image sequence includes a dynamic-static partition image corresponding to the target floor plan. The dynamic-static partition image includes multiple frames of dynamic-static partition images, and the dynamic-static partition marking states corresponding to different images are different. For example, the first frame image only marks the dynamic area, the second frame image only marks the static area, and the third frame image uses different marking styles to mark the dynamic and static areas respectively. Alternatively, the first frame image can use a first marking style (such as red) to mark the dynamic area, and the second frame image can use a second marking style (such as green) to mark the static area based on the marking of the previous frame image.

[0136] When generating the second video corresponding to the dynamic / static partition feature tags, the dynamic / static partition feature text corpus corresponding to the feature tags is added to the dynamic / static partition animation graph to generate the video frame sequence corresponding to the dynamic / static partition feature tags. Then, the video frame sequence corresponding to the dynamic / static partition feature tags is matched with the second narration audio to generate the second video corresponding to the dynamic / static partition feature tags. The dynamic / static partition feature text corpus is added to the dynamic / static partition animation graph as subtitle information. The second narration audio corresponding to the dynamic / static partition feature tags is matched with the dynamic / static partition feature text corpus; that is, the audio content of the second narration audio corresponding to the dynamic / static partition feature tags is the same as the text content of the dynamic / static partition feature text corpus.

[0137] The background of the dynamic map of the static and dynamic partitions can be a non-transparent overlay. When adding the dynamic and static partition feature text corpus as subtitle information to the dynamic map of the static and dynamic partitions, dynamic and static partition feature labels can also be added. The subtitle information can be added on the overlay or superimposed on the dynamic map of the static and dynamic partitions.

[0138] By adding the dynamic and static zoning feature text corpus to the dynamic and static zoning animation, users can easily understand the dynamic and static zoning features corresponding to the target floor plan while browsing the image, along with the subtitle information.

[0139] Optionally, when the feature label is the reasonable feature label of the movement path, the target image sequence includes a dynamic map of movement path markings corresponding to the target floor plan. The dynamic map of movement path markings includes multiple frames of movement path marking images, and the movement path marking states corresponding to different images can be different. For example, the first frame image only marks visitor movement paths, the second frame image only marks housework movement paths, and the third frame image only marks home movement paths; or, the first frame image only marks visitor movement paths, the second frame image continues to mark housework movement paths based on visitor movement paths with different marking styles, and the third frame image continues to mark home movement paths based on visitor movement paths and housework movement paths with different marking styles; or, consecutive frames of images can correspond to a single movement path, and the movement path marking degrees corresponding to different images can be different. For example, the first frame image, the second frame image, and the third frame image all correspond to visitor movement paths, but the first frame image only marks part of the movement path, the second frame image continues to mark based on the aforementioned movement path, and so on until the third frame image completes the marking. The three frames of images work together to form a dynamic effect of marking visitor movement paths.

[0140] When generating the second video corresponding to the reasonable movement feature tags, the reasonable movement feature text corpus corresponding to the reasonable movement feature tags is added to the movement mark animation graph to generate the video frame sequence corresponding to the reasonable movement feature tags. Then, the video frame sequence corresponding to the reasonable movement feature tags is matched with the second narration audio to generate the second video corresponding to the reasonable movement feature tags. The reasonable movement feature text corpus is added as subtitle information to the movement mark animation graph, and the second narration audio corresponding to the reasonable movement feature tags is matched with the reasonable movement feature text corpus; that is, the audio content of the second narration audio corresponding to the reasonable movement feature tags is the same as the text content of the reasonable movement feature text corpus.

[0141] The background of the motion-line marker animation can be a non-transparent overlay. When adding text corpus with reasonable motion-line features as subtitle information to the motion-line marker animation, reasonable motion-line feature tags can also be added. The subtitle information can be added on the overlay or superimposed on the motion-line marker animation. See also Figure 7 The image shows a specific example of adding circulation routes to a target floor plan.

[0142] By adding textual data with reasonable characteristics of the circulation path to the dynamic graph of the circulation path marker, users can better understand the circulation path distribution corresponding to the target floor plan while browsing the image, in conjunction with the subtitle information.

[0143] Optionally, when the feature label is the reasonable feature label for width and depth, the target image sequence includes a dynamic image of width and depth analysis corresponding to the target floor plan. The dynamic image of width and depth analysis may include a dynamic image of floor plan width and depth analysis and / or a dynamic image of real-scene width and depth analysis. The dynamic image of floor plan width and depth analysis includes multiple frames of floor plan images, and the analysis states corresponding to different images may be different. Here, the analysis state may be a marking state where the width and / or depth of certain rooms are marked. The dynamic image of real-scene width and depth analysis is generated by marking the width and depth after taking real-scene photos of the width and depth of a specific room.

[0144] When generating the second video corresponding to the reasonable feature tags of width and depth, the reasonable feature text corpus corresponding to the reasonable feature tags of width and depth is added to the width and depth analysis animation to generate the video frame sequence corresponding to the reasonable feature tags of width and depth. Then, the video frame sequence corresponding to the reasonable feature tags of width and depth is matched with the second narration audio to generate the second video corresponding to the reasonable feature tags of width and depth. Among them, the reasonable feature text corpus of width and depth is added as subtitle information to the width and depth analysis animation. The second narration audio corresponding to the reasonable feature tags of width and depth is adapted to the reasonable feature text corpus of width and depth, that is, the audio content of the second narration audio corresponding to the reasonable feature tags of width and depth is the same as the text content of the reasonable feature text corpus of width and depth.

[0145] By adding the text corpus of reasonable features of width and depth to the dynamic graph of width and depth analysis, users can easily understand the width and depth of the rooms in the target floor plan while browsing the image, along with the subtitle information.

[0146] The following section introduces the relationship between feature labels and the transparency and lighting dimension. When feature labels are associated with the transparency and lighting dimension, the feature labels may include at least one of the following: ventilation feature labels, apartment type feature labels, and reasonable lighting feature labels.

[0147] Optionally, when the feature label is the ventilation feature label, the target image sequence includes a two-dimensional ventilation effect animation and / or a three-dimensional ventilation effect animation corresponding to the target floor plan. The two-dimensional ventilation effect animation may include multiple frames of two-dimensional ventilation effect images, and the three-dimensional ventilation effect animation may include multiple frames of three-dimensional ventilation effect images. The ventilation effect labeling states corresponding to different images may differ. Furthermore, the two-dimensional and / or three-dimensional ventilation effect animations are adapted to the direct ventilation, two-sided ventilation, or one-sided ventilation indicated by the ventilation feature label.

[0148] When generating the second video corresponding to the ventilation feature tags, the ventilation feature text corpus corresponding to the ventilation feature tags is added to the 2D and / or 3D ventilation effect animations to generate a video frame sequence corresponding to the ventilation feature tags. Then, the video frame sequence corresponding to the ventilation feature tags is matched with the second narration audio to generate the second video corresponding to the ventilation feature tags. The ventilation feature text corpus is added as subtitle information to the 2D and / or 3D ventilation effect animations. The second narration audio corresponding to the ventilation feature tags is adapted to the ventilation feature text corpus; that is, the audio content of the second narration audio corresponding to the ventilation feature tags is the same as the text content of the ventilation feature text corpus.

[0149] See Figure 8a The image shown is a specific example of adding subtitle information to a two-dimensional ventilation rendering; see also... Figure 8b The image shown is a specific example of adding subtitle information to a 3D ventilation rendering.

[0150] By adding ventilation feature text corpus to two-dimensional and / or three-dimensional ventilation effect animations, users can easily understand the ventilation situation of the target floor plan while browsing the ventilation effect images, along with subtitle information.

[0151] Optionally, when the feature label is the apartment type feature label, the target image sequence includes a dynamic image of window markings corresponding to the target apartment layout. The dynamic image of window markings includes multiple frames of window marking images. The windows marked in different images may be different, or the window marking states corresponding to different images may be different. Here, the marking state can be a marking style. In this embodiment, the apartment type is related to the orientation of the windows in the room.

[0152] When generating the second video corresponding to the apartment type feature tags, the text corpus of the apartment type feature tags is added to the form marker animation to generate the video frame sequence corresponding to the apartment type feature tags. Then, the video frame sequence corresponding to the apartment type feature tags is matched with the second narration audio to generate the second video corresponding to the apartment type feature tags. The apartment type feature text corpus is added to the form marker animation as subtitle information. The second narration audio corresponding to the apartment type feature tags is matched with the apartment type feature text corpus; that is, the audio content of the second narration audio corresponding to the apartment type feature tags is the same as the text content of the apartment type feature text corpus.

[0153] See Figure 9 The image shown is a specific example of adding caption information to a form marker image.

[0154] By adding the text corpus of apartment type category characteristics to the form marker animation, users can easily understand the apartment type details while browsing the form marker animation, along with subtitle information.

[0155] Optionally, when the feature label is a reasonable lighting feature label, the target image sequence includes a dynamic image of the lighting effect corresponding to the target floor plan. This dynamic image may include a dynamic image of the floor plan's lighting effect and / or a dynamic image of the actual lighting effect. The dynamic image of the floor plan's lighting effect may include a two-dimensional lighting effect image and / or a three-dimensional lighting effect image. Specifically, the dynamic image of the floor plan's lighting effect includes multiple frames of lighting effect images. Since the dynamic image of the floor plan's lighting effect can include two-dimensional and / or three-dimensional lighting effect images, it is possible to obtain lighting effects in two-dimensional and / or three-dimensional scenes. The dynamic image of the actual lighting effect includes multiple frames of the actual lighting effect images.

[0156] When generating the second video corresponding to the "reasonable lighting feature tag," the text corpus corresponding to the "reasonable lighting feature tag" is added to the lighting effect animation to generate the video frame sequence corresponding to the "reasonable lighting feature tag." Then, the video frame sequence corresponding to the "reasonable lighting feature tag" is matched with the second narration audio to generate the second video corresponding to the "reasonable lighting feature tag." The "reasonable lighting feature text corpus" is added as subtitle information to the lighting effect animation. The second narration audio corresponding to the "reasonable lighting feature tag" is matched with the "reasonable lighting feature text corpus," meaning the audio content of the second narration audio corresponding to the "reasonable lighting feature tag" is the same as the text content of the "reasonable lighting feature text corpus."

[0157] See Figure 10 The image shown is a specific example of adding subtitle information to a two-dimensional lighting effect diagram.

[0158] By adding textual data on reasonable lighting characteristics to the dynamic lighting effect graph, users can better understand the lighting conditions of the target floor plan while browsing the lighting effect based on the image, along with subtitle information.

[0159] In one embodiment of this application, after generating the tour video, the method further includes: generating a decoration effect video based on the panoramic dynamic image of the decoration corresponding to the target floor plan; and stitching the tour video and the decoration effect video together to generate the target video.

[0160] In this embodiment, a panoramic dynamic image of the interior design corresponding to the target floor plan can be obtained. This panoramic dynamic image can be understood as a dynamic image of a design sample. Interior design corpus is added to the panoramic dynamic image, and then the dynamic image with added corpus is matched with the corresponding interior design explanation audio to generate an interior design effect video. After generating the interior design effect video, it is stitched together with a guided tour video. This video stitching generates a video that includes introductions to the floor plan's features and the interior design effect of the model unit, allowing users to understand the floor plan's features while also seeing the model interior design effect of the corresponding property based on the target floor plan.

[0161] The above describes the overall implementation process of the video generation method provided in this application. Based on the feature parameters of the target floor plan in the target dimension, N feature labels corresponding to the target floor plan are determined. A first video is generated to provide an overall introduction to the target floor plan based on the N feature labels, the target floor plan, a panoramic dynamic image of the community to which the target floor plan belongs, and a first explanatory audio. For each of the N feature labels, a second video is generated to individually introduce the floor plan features of that feature label based on the corresponding video frame sequence and the second explanatory audio. The first video and the N second videos are then stitched together to generate a guided tour video of the target floor plan. This method enables the generation of guided tour videos based on the floor plan features of the target floor plan, facilitating users to efficiently and accurately gain an overall understanding of the target floor plan and to understand the different features of the target floor plan based on each feature label. This allows users to gain a relatively comprehensive understanding of the floor plan features and improves their information acquisition experience.

[0162] Furthermore, based on video frame sequence splicing and audio-visual matching, a first video is generated to provide an overall introduction to the target floor plan, which can enable users to understand the target floor plan as a whole during subsequent presentations.

[0163] By adding the corresponding text corpus to the target image sequence for each feature label to generate a video frame sequence, and matching the video frame sequence with the narration audio to obtain the corresponding second video, users can easily understand the relevant information of the target floor plan while browsing the image, in conjunction with subtitle and audio information.

[0164] This application provides a video generation apparatus, see [link to relevant documentation]. Figure 11 As shown, it includes:

[0165] The determining module 1101 is used to determine N feature labels corresponding to the target floor plan based on the feature parameters of the target floor plan in the target dimension, where N is an integer greater than or equal to 1;

[0166] The first generation module 1102 is used to generate a first video based on the N feature tags, the target floor plan, the panoramic dynamic map of the community, and the first explanatory audio.

[0167] The second generation module 1103 is used to generate a second video corresponding to each feature label based on the video frame sequence corresponding to the feature label and the second narration audio.

[0168] The third generation module 1104 is used to combine the first video with N second videos to generate a tour video corresponding to the target floor plan.

[0169] Optionally, the target dimension includes at least one of the spatial scale dimension and the light transmittance dimension;

[0170] The N feature labels include at least one of the following associated with the spatial scale dimension: high space utilization feature label, square apartment layout feature label, structural feature label, active and quiet zone feature label, reasonable circulation feature label, and reasonable width and depth feature label, and / or, the N feature labels include at least one of the following associated with the transparency and lighting dimension: ventilation feature label, apartment type feature label, and reasonable lighting feature label.

[0171] Optionally, the first generation module includes:

[0172] The first generation submodule is used to generate a sequence of overall introduction video frames corresponding to the target floor plan based on the N feature labels, the target floor plan, the panoramic dynamic image of the community, and the first text corpus. The panoramic dynamic image of the community is the panoramic dynamic image corresponding to the target community to which the target floor plan belongs.

[0173] The second generation submodule is used to generate a spatially distinctive video frame sequence based on the target floor plan, the panoramic dynamic map of the community, and the second text corpus.

[0174] The third generation submodule is used to generate a sequence of labeled display video frames based on the target floor plan and the N feature labels;

[0175] The fourth generation submodule is used to generate a sequence of video frames for displaying the floor plan based on the target floor plan, motion effect rules, and third text corpus.

[0176] The fifth generation submodule is used to combine the overall introduction video frame sequence, the spatial feature video frame sequence, the tag display video frame sequence, and the floor plan display video frame sequence to generate the first video frame sequence;

[0177] The sixth generation submodule is used to match the first video frame sequence and the first narration audio to generate the first video. The text corpus corresponding to the first narration audio includes the first text corpus, the second text corpus, the tag text corpus matching the N feature tags, and the third text corpus.

[0178] Optionally, the first generation submodule includes one of the following units:

[0179] The first generation unit is used to add the N feature labels and the first text corpus to the panoramic dynamic map of the community to generate a first target dynamic map, and to stitch each frame of the first target dynamic map with the target floor plan to generate the overall introduction video frame sequence.

[0180] The second generation unit is used to stitch together each frame of the panoramic dynamic map of the community with the target floor plan to generate a second target dynamic map, and to add the N feature labels and the first text corpus to the panoramic area of ​​the community in the second target dynamic map to generate the overall introduction video frame sequence.

[0181] The duration of the overall introduction video frame sequence is equal to the duration of the first text corpus.

[0182] Optionally, the second generation submodule is further used for:

[0183] Using the panoramic dynamic image of the community as the background and the target apartment floor plan as the main body, a sequence of spatially characteristic video frames with the second text corpus added is generated.

[0184] The duration of the spatially characteristic video frame sequence is equal to the duration of the second text corpus.

[0185] Optionally, the third generation submodule is further configured to:

[0186] The N feature labels are sequentially added to the target floor plan to generate the label display video frame sequence;

[0187] Wherein, the duration of the video frame sequence displayed by the tag is equal to the duration of the tag text corpus.

[0188] Optionally, the fourth generation submodule is further configured to:

[0189] Generate an animated floor plan based on the animation rules and the target floor plan;

[0190] The third text corpus is added to the animated floor plan to generate a sequence of video frames displaying the floor plan.

[0191] The duration of the video frame sequence displaying the floor plan is equal to the duration of the third text corpus.

[0192] Optionally, when the value of N is greater than or equal to 2, the third generation module includes:

[0193] The sorting submodule is used to sort the N second videos according to the priority of the feature tags;

[0194] The video splicing generation submodule is used to splice the first video and the sorted N second videos to generate the tour video.

[0195] Optionally, the device further includes:

[0196] The fourth generation module is used to generate a decoration effect video based on the decoration panoramic dynamic map corresponding to the target floor plan after the guide video is generated by the third generation module.

[0197] The fifth generation module is used to stitch the tour video and the decoration effect video together to generate the target video.

[0198] Optionally, the second generation module includes:

[0199] A generation submodule is added to add the text corpus corresponding to the feature label to the target image sequence corresponding to the feature label, thereby generating the video frame sequence corresponding to the feature label.

[0200] The matching and generation submodule is used to match the video frame sequence corresponding to the feature label and the second narration audio to generate the second video corresponding to the feature label.

[0201] The second audio explanation corresponding to the feature label is adapted to the text corpus corresponding to the feature label.

[0202] Optionally, when the feature labels are associated with the spatial scale dimension, the target image sequence is as follows:

[0203] When the feature label is the high space utilization feature label, the target image sequence includes the outline marker dynamic image and / or real-scene dynamic image corresponding to the target floor plan;

[0204] When the feature label is the square floor plan feature label, the target image sequence includes the target floor plan;

[0205] When the feature label is the structural feature label, the target image sequence includes a structural analysis animation and / or a structural real-scene animation corresponding to the structural feature label;

[0206] When the feature label is the dynamic and static partition feature label, the target image sequence includes the dynamic and static partition dynamic map corresponding to the target floor plan;

[0207] When the feature label is the reasonable feature label of the movement path, the target image sequence includes a dynamic map of the movement path markings corresponding to the target floor plan;

[0208] When the feature label is the reasonable feature label of the width and depth, the target image sequence includes a dynamic image of the width and depth analysis corresponding to the target floor plan.

[0209] Optionally, when the feature label is associated with the light transmittance dimension, the target image sequence is as follows:

[0210] When the feature label is the ventilation feature label, the target image sequence includes a two-dimensional ventilation effect animation and / or a three-dimensional ventilation effect animation corresponding to the target floor plan;

[0211] When the feature label is the apartment type feature label, the target image sequence includes a dynamic window marker image corresponding to the target apartment layout.

[0212] When the feature label is the light-reasoning feature label, the target image sequence includes a dynamic image of the light-reasoning effect corresponding to the target floor plan.

[0213] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0214] This application also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the various processes of the above-described video generation method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0215] For example, Figure 12 A schematic diagram of the physical structure of an electronic device is shown. (For example...) Figure 12 As shown, the electronic device may include a processor 1210, a communications interface 1220, a memory 1230, and a communication bus 1240, wherein the processor 1210, the communications interface 1220, and the memory 1230 communicate with each other via the communication bus 1240. The processor 1210 can call logical instructions in the memory 1230 to perform the following steps: determining N feature labels corresponding to the target floor plan based on feature parameters of the target floor plan in the target dimension, where N is an integer greater than or equal to 1; generating a first video based on the N feature labels, the target floor plan, a panoramic dynamic image of the community, and a first explanatory audio; generating a second video corresponding to each feature label based on the video frame sequence corresponding to the feature label and the second explanatory audio; and combining the first video with the N second videos to generate a guided tour video corresponding to the target floor plan. The processor 1210 may also execute other schemes in the embodiments of this application, which will not be further described here.

[0216] Furthermore, the logical instructions in the aforementioned memory 1230 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0217] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described video generation method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0218] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0219] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0220] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0221] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0222] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0223] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0224] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0225] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0226] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0227] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A video generation method, characterized in that, include: Based on the feature parameters of the target floor plan in the target dimension, determine N feature labels corresponding to the target floor plan, where N is an integer greater than or equal to 1; the target dimension includes at least one of the spatial scale dimension and the transparency and lighting dimension. Specifically, when generating feature labels based on the feature parameters of the spatial scale dimension, a high space utilization feature label is generated when the space utilization rate corresponding to the target floor plan meets the first preset condition; a square floor plan feature label is generated when the squareness of the floor plan corresponding to the target floor plan meets the second preset condition; a dynamic and static zoning feature label is generated when the room distribution of the target floor plan meets the dynamic and static zoning; and a reasonable circulation feature label is generated when the room distribution of the target floor plan meets the reasonable circulation layout. When generating feature labels based on the characteristic parameters of the transparency and light transmission dimension, ventilation feature labels are generated according to the ventilation conditions. Specifically, if the ventilation conditions indicate that the ventilation type is direct ventilation, a direct ventilation feature label is generated; if the ventilation conditions indicate that the ventilation type is two-sided ventilation, a two-sided ventilation feature label is generated; if the ventilation conditions indicate that the ventilation type is one-sided ventilation, a one-sided ventilation feature label is generated; and a unit type feature label is generated according to the unit type. A first video is generated based on the N feature tags, the target floor plan, the panoramic dynamic image of the community, and the first explanatory audio. For each feature tag, a second video corresponding to the feature tag is generated based on the video frame sequence corresponding to the feature tag and the second narration audio. The first video is combined with N second videos to generate a guided video corresponding to the target floor plan; The step of generating a first video based on the N feature tags, the target apartment floor plan, the panoramic dynamic image of the community, and the first explanatory audio includes: Based on the N feature labels, the target floor plan, the community panoramic dynamic image, and the first text corpus, a sequence of overall introduction video frames corresponding to the target floor plan is generated, and the community panoramic dynamic image is the panoramic dynamic image corresponding to the target community to which the target floor plan belongs. Based on the target floor plan, the panoramic dynamic image of the community, and the second text corpus, a spatially distinctive video frame sequence is generated; Based on the target floor plan and the N feature labels, generate a sequence of labeled display video frames; Based on the target floor plan, motion effect rules, and third-party text corpus, generate a sequence of video frames for displaying the floor plan. The first video frame sequence is generated by combining the overall introduction video frame sequence, the spatial feature video frame sequence, the tag display video frame sequence, and the floor plan display video frame sequence. The first video frame sequence and the first narration audio are matched to generate the first video. The text corpus corresponding to the first narration audio includes the first text corpus, the second text corpus, the tag text corpus that matches the N feature tags, and the third text corpus.

2. The method according to claim 1, characterized in that, The step of generating a sequence of overall introduction video frames corresponding to the target floor plan based on the N feature labels, the target floor plan, the panoramic dynamic image of the community, and the first text corpus includes one of the following schemes: The N feature labels and the first text corpus are added to the panoramic dynamic map of the community to generate the first target dynamic map. Each frame of the first target dynamic map is stitched together with the target floor plan to generate the overall introduction video frame sequence. Each frame of the panoramic dynamic image of the community is stitched together with the target floor plan to generate a second target dynamic image. The N feature labels and the first text corpus are added to the panoramic area of ​​the community in the second target dynamic image to generate the overall introduction video frame sequence. The duration of the overall introduction video frame sequence is equal to the duration of the first text corpus.

3. The method according to claim 1, characterized in that, The step of generating a spatially characteristic video frame sequence based on the target apartment floor plan, the panoramic dynamic image of the community, and the second text corpus includes: Using the panoramic dynamic image of the community as the background and the target apartment floor plan as the main body, a sequence of spatially characteristic video frames with the second text corpus added is generated. The duration of the spatially characteristic video frame sequence is equal to the duration of the second text corpus.

4. The method according to claim 1, characterized in that, The step of generating a sequence of labeled display video frames based on the target floor plan and the N feature labels includes: The N feature labels are sequentially added to the target floor plan to generate the label display video frame sequence; Wherein, the duration of the video frame sequence displayed by the tag is equal to the duration of the tag text corpus.

5. The method according to claim 1, characterized in that, The step of generating a sequence of video frames for displaying the floor plan based on the target floor plan, motion effect rules, and third-party text corpus includes: Generate an animated floor plan based on the animation rules and the target floor plan; The third text corpus is added to the animated floor plan to generate a sequence of video frames displaying the floor plan. The duration of the video frame sequence displaying the floor plan is equal to the duration of the third text corpus.

6. The method according to claim 1, characterized in that, When N is greater than or equal to 2, the step of combining the first video with N second videos to generate a guided video corresponding to the target floor plan includes: Sort the N second videos according to the priority of their feature tags; The first video and the N sorted second videos are spliced ​​together to generate the tour video.

7. The method according to claim 1, characterized in that, After generating the guided tour video, the following is also included: Generate a renovation effect video based on the panoramic dynamic image of the renovation corresponding to the target floor plan; The guide video and the decoration effect video are spliced ​​together to generate the target video.

8. The method according to claim 1, characterized in that, The step of generating a second video corresponding to the feature label based on the video frame sequence corresponding to the feature label and the second narration audio includes: The text corpus corresponding to the feature label is added to the target image sequence corresponding to the feature label to generate the video frame sequence corresponding to the feature label; The video frame sequence corresponding to the feature label and the second audio explanation are matched to generate the second video corresponding to the feature label. The second audio explanation corresponding to the feature label is adapted to the text corpus corresponding to the feature label.

9. The method according to claim 8, characterized in that, When the feature labels are associated with the spatial scale dimension, the target image sequence is as follows: When the feature label is the high space utilization feature label, the target image sequence includes the outline marker dynamic image and / or real-scene dynamic image corresponding to the target floor plan; When the feature label is the square floor plan feature label, the target image sequence includes the target floor plan; When the feature label is a structural feature label, the target image sequence includes a structural analysis animation and / or a structural real-scene animation corresponding to the structural feature label; When the feature label is the dynamic and static partition feature label, the target image sequence includes the dynamic and static partition dynamic map corresponding to the target floor plan; When the feature label is the reasonable feature label of the movement path, the target image sequence includes a dynamic map of the movement path markings corresponding to the target floor plan; When the feature label is a reasonable feature label for width and depth, the target image sequence includes a dynamic analysis diagram of width and depth corresponding to the target floor plan.

10. The method according to claim 8, characterized in that, When the feature label is associated with the light transmittance dimension, the target image sequence is as follows: When the feature label is the ventilation feature label, the target image sequence includes a two-dimensional ventilation effect animation and / or a three-dimensional ventilation effect animation corresponding to the target floor plan; When the feature label is the apartment type feature label, the target image sequence includes a dynamic window marker image corresponding to the target apartment layout. When the feature label is a reasonable lighting feature label, the target image sequence includes a dynamic image of the lighting effect corresponding to the target floor plan.

11. A video generation apparatus, characterized in that, include: The determination module is used to determine N feature labels corresponding to the target floor plan based on the feature parameters of the target floor plan in the target dimension, where N is an integer greater than or equal to 1; the target dimension includes at least one of spatial scale dimension and light transmission dimension; wherein, when generating feature labels based on the feature parameters of the spatial scale dimension, a high space utilization feature label is generated when the space utilization of the target floor plan meets a first preset condition; a square floor plan feature label is generated when the squareness of the floor plan meets a second preset condition; and a square floor plan feature label is generated based on the room distribution of the target floor plan. When the separation of active and quiet zones is satisfied, a feature label for active and quiet zone separation is generated; when the room distribution in the target floor plan satisfies the requirement of reasonable circulation layout, a feature label for reasonable circulation is generated; when generating feature labels based on the feature parameters of the transparency and lighting dimension, a ventilation feature label is generated based on the ventilation status, where if the ventilation status indicates that the ventilation type is direct ventilation, a direct ventilation feature label is generated; if the ventilation status indicates that the ventilation type is two-sided ventilation, a two-sided ventilation feature label is generated; if the ventilation status indicates that the ventilation type is one-sided ventilation, a one-sided ventilation feature label is generated; and a floor plan type feature label is generated based on the floor plan type. The first generation module is used to generate a first video based on the N feature tags, the target floor plan, the panoramic dynamic map of the community, and the first explanatory audio; the second generation module is used to generate a second video corresponding to each feature tag based on the video frame sequence corresponding to the feature tag and the second explanatory audio. The third generation module is used to combine the first video with N second videos to generate a tour video corresponding to the target floor plan; The first generation module includes: The first generation submodule is used to generate a sequence of overall introduction video frames corresponding to the target floor plan based on the N feature labels, the target floor plan, the panoramic dynamic image of the community, and the first text corpus. The panoramic dynamic image of the community is the panoramic dynamic image corresponding to the target community to which the target floor plan belongs. The second generation submodule is used to generate a spatially distinctive video frame sequence based on the target floor plan, the panoramic dynamic map of the community, and the second text corpus. The third generation submodule is used to generate a sequence of labeled display video frames based on the target floor plan and the N feature labels; The fourth generation submodule is used to generate a sequence of video frames for displaying the floor plan based on the target floor plan, motion effect rules, and third text corpus. The fifth generation submodule is used to combine the overall introduction video frame sequence, the spatial feature video frame sequence, the tag display video frame sequence, and the floor plan display video frame sequence to generate the first video frame sequence; The sixth generation submodule is used to match the first video frame sequence and the first narration audio to generate the first video. The text corpus corresponding to the first narration audio includes the first text corpus, the second text corpus, the tag text corpus matching the N feature tags, and the third text corpus.

12. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the video generation method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the video generation method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and device for automatically generating house explanation in virtual three-dimensional space

    CN110110104A

  • House resource information processing method and device

    CN112596694A

  • Image processing method and device, electronic equipment and storage medium

    CN114827574A