Multi-modal AI-based house property house exploration video full-process intelligent generation system
By automatically acquiring property features and lighting conditions through a multimodal AI system, high-quality house visit videos are generated, solving the problems of low efficiency and poor quality in existing technologies and achieving efficient and professional video production.
Patent Information
- Application Number
- CN202610228225.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-26
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies for producing house visit videos are inefficient and of poor quality, making it difficult to meet the needs of large-scale, multi-house video generation, and the quality of the finished products varies.
A full-process intelligent generation system for real estate exploration videos based on multimodal AI is adopted. The system acquires real estate features through a feature acquisition module, matches lighting conditions through a condition acquisition module, generates exploration videos through a video editing module, and performs quality verification and optimization through an editing optimization module, thereby achieving automated editing and multi-dimensional quality verification.
It significantly improves the automation and efficiency of house visit videos, ensuring professionalism and visual comfort in terms of lighting, smooth movement, and key point display, and has the ability to continuously iterate to improve the overall video quality.
Smart Images

Figure CN122069414A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video production optimization technology, specifically to a full-process intelligent generation system for real estate exploration videos based on multimodal AI. Background Technology
[0002] With the rapid development of digital marketing in the real estate industry, property tour videos have become a core medium for understanding property information and experiencing the space. High-quality property tour videos not only need to comprehensively showcase the key highlights of the property, but also need to have good visual appeal and efficient information delivery in terms of lighting, smooth movement, and other aspects. Currently, on the one hand, manual production methods are time-consuming, labor-intensive, and costly, making it difficult to meet the needs of large-scale, multi-property video generation, and the quality of the finished products varies greatly. Summary of the Invention
[0003] This application provides a full-process intelligent generation system for real estate visit videos based on multimodal AI, which is used to address the technical problems of low efficiency and poor quality in the production of real estate visit videos in the prior art.
[0004] In view of the above problems, this application provides a full-process intelligent generation system for real estate exploration videos based on multimodal AI.
[0005] In a first aspect, this application provides an intelligent generation system for the entire process of real estate exploration videos based on multimodal AI, the system comprising:
[0006] The feature acquisition module is used to acquire the property features of the property to be generated in the video.
[0007] The condition acquisition module is used to obtain the functional area weights based on the property features, and to perform lighting matching in combination with the property features to obtain the lighting conditions.
[0008] The video editing module is used to obtain the time sequence of the functional area based on the functional area weight and the lighting conditions, form a filtering constraint set and an editing constraint set, and filter the materials to edit and generate a house visit video.
[0009] The editing optimization module is used to perform quality verification on the house visit video, obtain editing adaptation parameters, optimize the filtering constraint set and the editing constraint set, obtain the optimized constraint set, and filter materials to generate an optimized house visit video.
[0010] Secondly, this application provides a method for intelligent generation of real estate exploration videos throughout the entire process based on multimodal AI, including:
[0011] Obtain the property features of the property to be used in the video;
[0012] Based on the property characteristics, functional area weights are obtained, and lighting conditions are obtained by combining the property characteristics with lighting matching.
[0013] Based on the functional area weights and the lighting conditions, the time sequence of the functional areas is obtained, forming a set of filtering constraints and a set of editing constraints. The selected materials are then edited to generate a house visit video.
[0014] The quality of the house visit video is verified, editing adaptation parameters are obtained, the filtering constraint set and the editing constraint set are optimized, an optimized constraint set is obtained, and the selected materials are used to generate an optimized house visit video.
[0015] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0016] This application proposes a multimodal AI-based intelligent generation system for the entire process of real estate exploration videos. By deeply integrating multi-dimensional real estate features with intelligent content production processes, it constructs a complete solution for generating real estate exploration videos, encompassing feature analysis, lighting matching, material selection, automated editing, quality verification, and closed-loop optimization. The system first uses a machine learning-trained importance allocation model to intelligently deconstruct the relative importance of different functional areas based on the real estate features to be generated, including functional area features, orientation features, and key features. This model generates normalized functional area weights, providing objective and quantifiable decision-making basis for subsequent content arrangement. Simultaneously, by combining the correspondence between orientation features and natural lighting periods, it automatically matches the optimal lighting angle and shooting time, fundamentally ensuring the professionalism and visual comfort of the video's lighting presentation. Building upon this foundation, the system adaptively generates presentation priority sequences, duration requirements, and filtering and editing constraint sets for functional areas based on functional area weights and lighting conditions. This enables the system to accurately select highly relevant content from source material and automatically edit and arrange it according to professional narrative logic, resulting in highly coordinated room visit videos in terms of functional display completeness, smooth movement, and emphasis on key points. Furthermore, the system performs multi-dimensional quality checks on the initial room visit video, comprehensively calculating editing adaptation parameters from the perspectives of priority matching, spatial movement smoothness, and lighting uniformity. Based on these parameters, the system dynamically optimizes the filtering and editing constraint sets. Through a deviation-driven optimization set generation and constraint set fine-tuning mechanism, the system selects and outputs the highest-quality optimized room visit video from multiple candidate videos. Compared to traditional methods, the technical solution provided in this application significantly improves the automation and efficiency of room visit video production. In addition, the closed-loop quality check and constraint set optimization enables the system to continuously iterate, dynamically adjusting the generation strategy based on the quality feedback of the final video, significantly enhancing the overall quality of the output video. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of the structure of the intelligent generation system for the entire process of real estate exploration video based on multimodal AI, which is provided in the embodiments of this application.
[0019] Figure 2 This is a flowchart illustrating the intelligent generation method for the entire process of property exploration videos based on multimodal AI, provided in an embodiment of this application.
[0020] The components represented by each number in the attached diagram are explained below:
[0021] Feature acquisition module 100, condition acquisition module 200, video editing module 300, and editing optimization module 400. Detailed Implementation
[0022] This application provides a multimodal AI-based intelligent generation system for the entire process of real estate exploration videos, which aims to address the technical problems of low efficiency and poor quality in existing real estate exploration video production.
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0024] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0025] Example 1, as Figure 1 As shown, this application provides a multimodal AI-based intelligent generation system for the entire process of real estate exploration videos, wherein the system includes:
[0026] The feature acquisition module 100 is used to acquire the property features of the property to be generated in the video.
[0027] In the process of intelligently generating real estate exploration videos, the logical starting point for realizing all subsequent personalized decisions and content arrangement is to enable the system to truly understand and accurately represent the core characteristics of a property to be filmed.
[0028] In the system provided in this application embodiment, the execution flow of the feature acquisition module 100 includes:
[0029] Obtain the property features of the property to be used in the video, wherein the property features include functional area features, orientation features, and key features.
[0030] Specifically, in this embodiment, the property features of the property for which the video is to be generated are obtained. These features include functional area features, orientation features, and key features. For example, the living room area is approximately 30 square meters, the master bedroom area is approximately 20 square meters, the kitchen area is approximately 8 square meters, and the bathroom area is approximately 5 square meters, including spatial location data for each functional area. Further, orientation features are extracted from the dataset. Orientation features include the overall orientation of the property's main lighting surfaces, as well as the orientation of the windows or balconies of each main functional area. For example, it is obtained that the living room and master bedroom face south, the secondary bedroom faces north, and the kitchen faces east. Further, key features are obtained from the dataset. Key features are derived from preset labels or manually labeled key points; for example, key features of the property include "floor-to-ceiling balcony" and "walk-in closet."
[0031] The feature acquisition module enables the systematic and structured collection and representation of property features in the video to be generated, providing accurate input for the subsequent importance allocation model.
[0032] The condition acquisition module 200 is used to acquire functional area weights based on the property features, and to perform lighting matching in combination with the property features to acquire lighting conditions.
[0033] Different properties naturally have different functional areas, and current technology lacks objective and reusable quantitative methods to adaptively allocate the relative importance of each functional area based on the characteristics of the property. This results in the exposure time and display order of key areas in videos not matching the actual attention of users. In addition, lighting conditions are a key factor determining the quality and professionalism of video images. Functional areas with different orientations present drastically different lighting effects at different times, resulting in poor visual performance in some areas.
[0034] In the system provided in this application embodiment, the execution flow of the condition acquisition module 200 includes:
[0035] The property features are input into an importance allocation model to obtain functional area importance parameters;
[0036] The configuration of the importance allocation model includes:
[0037] Historical house visit videos were acquired, edited, and analyzed to obtain functional area importance parameters, which were then used as a sample output set.
[0038] Collect the property characteristics of the properties corresponding to the historical house visit videos, and use them as the sample input set;
[0039] Based on machine learning, an importance allocation model is constructed, and the importance allocation model is trained using the sample input set and the sample output set until convergence, thus completing the configuration of the importance allocation model;
[0040] The importance parameters of the multiple functional areas are normalized to obtain the functional area weights;
[0041] By combining the orientation characteristics of the property, the optimal lighting angle can be obtained through analysis.
[0042] Based on the optimal lighting angle and the corresponding relationship of natural lighting time periods, the lighting conditions are derived, wherein the lighting conditions include the shooting time.
[0043] In this embodiment of the application, functional area weights are obtained based on the property characteristics, and lighting conditions are obtained by combining the property characteristics with lighting matching.
[0044] Specifically, firstly, the property characteristics are input into the importance allocation model to obtain the functional area importance parameters.
[0045] The configuration of the importance allocation model includes:
[0046] Historical property visit videos are acquired and edited for analysis to obtain functional area importance parameters, which serve as the sample output set. For example, a large number of historical property visit videos are retrieved from a database. The editing structure of each video is analyzed, and by statistically analyzing the actual on-screen time percentage of each functional area, the frequency of mention in the narration, and the priority position of each area in the final cut, the true importance parameters of each functional area of the property are deduced, and this is used as the sample output set.
[0047] Furthermore, the property characteristics of the properties corresponding to the historical house visit videos are collected as a sample input set. For example, the complete property characteristics of the properties corresponding to these historical house visit videos, including functional area characteristics, orientation characteristics, and key characteristics, are collected as a sample input set.
[0048] Furthermore, based on machine learning, an importance allocation model is constructed, and the model is trained using the sample input set and the sample output set until convergence, thus completing the configuration of the importance allocation model. For example, it is constructed using the TensorFlow framework. The model structure consists of three layers: the first layer is the input layer, with the number of nodes equal to the dimension of the property feature vector, and each node corresponds to a feature (such as functional area area, orientation attribute, whether it is a key area, etc.); the second layer is the hidden layer, containing 128 neurons, using the ReLU activation function, and adding a Dropout layer to prevent overfitting; the third layer is the output layer, with the number of nodes equal to the total number of all possible functional area categories, using a linear activation function, and the value of each output node represents the original importance score of the corresponding functional area.
[0049] Furthermore, the property features are input into an importance allocation model to obtain functional area importance parameters. For example, the property feature vector of the property to be generated in the video is input into the model, the model performs forward propagation calculation, and outputs a set of original importance parameters, such as a score of 85 for the living room, 72 for the master bedroom, 60 for the kitchen, and 45 for the bathroom.
[0050] Furthermore, the importance parameters of the multiple functional areas are normalized to obtain the functional area weights. Specifically, the weight of a single functional area is equal to its original importance parameter divided by the sum of the original importance parameters of all functional areas. For example, if the sum of the original importance parameters of all functional areas is 85 + 72 + 60 + 45 = 262, then the weight of the living room is 85 ÷ 262 = 0.324, the weight of the master bedroom is 72 ÷ 262 = 0.275, the weight of the kitchen is 60 ÷ 262 = 0.229, and the weight of the bathroom is 45 ÷ 262 = 0.172. This set of normalized values is the final weight of each functional area, and its sum is 1, reflecting the relative importance proportion that each functional area should occupy in the house visit video.
[0051] Furthermore, by combining the orientation characteristics of the property features, the optimal lighting angle is analyzed and obtained. For example, the orientation information of each main functional area is extracted from the property features, such as the living room facing south, the master bedroom facing south, and the kitchen facing east. According to the basic principles of building lighting, south-facing spaces receive the best direct sunlight around noon, while east-facing spaces have the best lighting quality in the morning.
[0052] Furthermore, based on the optimal lighting angle and the corresponding relationship of natural lighting periods, lighting conditions are derived, including the shooting time. Specifically, for each functional area, the natural light intensity variation curve corresponding to its orientation throughout the day is analyzed, and the time window with uniform lighting, no strong shadows, and suitable light color temperature is selected as the time period corresponding to the optimal lighting angle of that functional area. Combining the optimal shooting times of all functional areas, the intersection window that can cover the most high-weight functional areas is selected to derive the overall shooting time. For example, if the living room (weight 0.324) and master bedroom (weight 0.275) are both south-facing and the optimal time is from 10:00 to 14:00, and the kitchen (weight 0.229) is east-facing and the optimal time is from 8:00 to 11:00, then after weighing the factors, 10:00 to 11:00 is selected as the starting shooting time. At this time, the south-facing space has sufficient light, and the east-facing space is still within the morning's excellent light period.
[0053] The condition acquisition module enables a crucial leap from property characteristics to video creation decision parameters. Its technological effectiveness is concentrated in two aspects: the quantitative allocation of functional area importance and the intelligent matching of lighting conditions. It can automatically plan a shooting window for each property without human intervention, avoiding strong shadows while capturing optimal lighting quality, significantly improving video image quality.
[0054] The video editing module 300 is used to obtain the time sequence of the functional area based on the functional area weight and the lighting conditions, form a filtering constraint set and an editing constraint set, and filter the materials to edit and generate a house visit video.
[0055] Existing automated editing tools are generally based on template splicing with fixed scripts, assembling preset segments in a fixed order, lacking reasonable planning for professional editing elements such as presentation logic, duration allocation, and flow continuity.
[0056] In the system provided in this application embodiment, the execution flow of the video editing module 300 includes:
[0057] Based on the functional area weights, obtain the presentation priority sequence and duration requirements for each functional area;
[0058] Based on the lighting conditions and duration requirements, obtain the set of filtering constraints;
[0059] Based on the presentation priority sequence and the duration requirement, obtain the editing constraint set;
[0060] The materials are filtered based on the filtering constraint set, and the filtered materials are edited in conjunction with the editing constraint set to generate a house visit video.
[0061] In this embodiment of the application, based on the functional area weights and the lighting conditions, the time sequence of the functional areas is obtained to form a filtering constraint set and an editing constraint set, and the selected materials are edited to generate a house visit video.
[0062] Specifically, firstly, based on the functional area weights, the presentation priority sequence and duration requirements for each functional area are obtained. For example, a list of weight values for each functional area is received from the condition acquisition module, such as a weight of 0.324 for the living room, 0.275 for the master bedroom, 0.229 for the kitchen, and 0.172 for the bathroom. A sorting function is used to sort the functional areas from highest to lowest weight, resulting in the presentation priority sequence: living room, master bedroom, kitchen, bathroom. Simultaneously, based on the preset total duration of the house visit video, the total duration is proportionally allocated according to the weights of each functional area, and the duration requirement for each functional area in the final video is calculated. For example, if the total duration is a certain value, the duration requirement for the living room is equal to that value multiplied by 0.324, the duration requirement for the master bedroom is equal to that value multiplied by 0.275, and so on, resulting in a list of duration requirements for each functional area.
[0063] Furthermore, based on the lighting conditions and duration requirements, a set of filtering constraints is obtained. For example, lighting conditions are received from the condition acquisition module, which include a clear description of the shooting time and optimal lighting angle, such as the shooting time being 10:00 AM to 11:00 AM. Based on this, the first constraint for material filtering is constructed: only video clips in the material library whose shooting time metadata matches this time period and whose shooting angle conforms to the optimal lighting angle are selected. Simultaneously, a second constraint for filtering is constructed by combining the duration requirements of each functional area: the total duration of the material clips required for each functional area should not be less than the duration requirement of that functional area, and the duration of a single clip should be suitable for editing and splicing. These constraints are encoded into a structured set of filtering constraints, including time conditions, angle conditions, duration conditions, and clip quantity conditions, for subsequent execution by the material retrieval module.
[0064] Furthermore, based on the presentation priority sequence and the duration requirement, an editing constraint set is obtained. For example, the presentation priority sequence is used as the core rule for arranging the video segments, constraining the order in which functional areas appear in the final cut to follow this sequence. The duration requirement of each functional area is used as the editing benchmark for retaining the segment length, stipulating that the duration extracted from the original footage selected from each functional area should be as close as possible to the duration requirement of that functional area. In addition, the editing constraint set also includes smooth transition requirements at segment connections, such as the physical distance between functional areas of adjacent segments being less than the maximum functional area distance. The above rules regarding order, duration, and transitions are uniformly encapsulated into an editing constraint set, serving as the execution instructions for subsequent automated editing.
[0065] Furthermore, the materials are filtered based on the filtering constraint set, and the filtered materials are edited in conjunction with the editing constraint set to generate a house visit video. By executing a database query, the conditions in the filtering constraint set are matched with the segment metadata. For example, all video segments tagged "living room," filmed in the morning, filmed from a south-facing angle with direct sunlight, and with an original duration greater than a certain lower limit are filtered out. For each functional area, one or more candidate segments that meet the conditions are retrieved, and these lists are output as the filtered materials. Further, the filtered materials are automatically edited in conjunction with the editing constraint set to generate the house visit video. For example, the open-source video editing engine FFmpeg can be called to arrange the candidate segments of each functional area according to the order requirements in the editing constraint set, based on their presentation priority. Within each functional area, the most suitable segment is selected from the candidate segment list according to the duration requirement, and its starting part is trimmed to the required duration using a cut command. All segments are spliced together to create a single, complete video file.
[0066] First, this module calculates the presentation priority sequence and duration requirements for each functional area based on its weight. Simultaneously, it constructs a set of filtering constraints based on lighting conditions and duration requirements, and then constructs an editing constraint set based on the priority sequence and duration requirements. This maps abstract property characteristics and lighting parameters into a set of clearly defined, quantifiable, and executable video production rules. Second, relying on the filtering constraint set, the module performs multi-dimensional matching and filtering of original video clips in the media library, automatically selecting high-quality footage that meets the current property requirements in terms of content theme, shooting angle, and lighting quality. This effectively avoids the efficiency bottlenecks and subjective biases of manual material selection. Finally, the module combines the editing constraint set to automatically edit the selected footage, strictly organizing the appearance order of functional areas according to the priority sequence, controlling the length of each segment according to the duration requirements, and incorporating basic transition and pacing logic. Ultimately, this generates a logically coherent, focused, and visually smooth complete property exploration video.
[0067] The editing optimization module 400 is used to perform quality verification on the house visit video, obtain editing adaptation parameters, optimize the filtering constraint set and the editing constraint set, obtain the optimized constraint set, and filter materials to generate an optimized house visit video.
[0068] A single-generation video often struggles to achieve optimal quality across multiple dimensions simultaneously. Existing technologies lack objective and quantifiable evaluation mechanisms for final video quality. Even when quality deficiencies are identified, there is a lack of effective feedback and optimization paths, failing to translate evaluation conclusions into corrective instructions for front-end editing strategies, resulting in poor video quality.
[0069] In the system provided in this application embodiment, the execution flow of the editing optimization module 400 includes:
[0070] Extract the presentation order of the functional areas in the house visit video, calculate the similarity with the presentation priority sequence, and obtain the priority parameter;
[0071] Based on the presentation order and physical location of the functional areas, obtain the fluency parameters;
[0072] Based on the illumination area sequence of the house visit video, brightness parameters are obtained;
[0073] The priority parameter, the smoothness parameter, and the brightness parameter are weighted and calculated to obtain the editing adaptation parameter;
[0074] Based on the editing adaptation parameters, obtain the number of optimized sets Q;
[0075] Among them, obtaining the number Q of optimized sets based on the editing adaptation parameters includes:
[0076] Calculate the deviation between the clip adaptation parameters and the clip adaptation threshold;
[0077] The number of optimization sets Q is obtained by multiplying the deviation by the preset number of optimization sets;
[0078] Based on the number of optimization sets, multiple material screenings and editing processes are performed to obtain multiple candidate house visit videos, and the candidate house visit video with the largest editing adaptation parameter is selected as the optimized house visit video.
[0079] Based on the number of optimized sets, multiple rounds of material screening and editing are performed to obtain multiple candidate house visit videos, including:
[0080] Fine-tuning the editing constraint set and the filtering constraint set generates Q optimized editing constraint sets and optimized filtering constraint sets, wherein the fine-tuning range is obtained based on the deviation degree;
[0081] The Q sets of optimized clipping constraints and optimized filtering constraints are randomly paired to obtain multiple constraint set pairs, wherein each constraint set pair includes an optimized clipping constraint set and an optimized filtering constraint set;
[0082] Multiple material screenings and editing processes were performed using the aforementioned constraint set to obtain multiple candidate house visit videos.
[0083] In this embodiment of the application, the quality of the house visit video is verified, editing adaptation parameters are obtained, the filtering constraint set and the editing constraint set are optimized, an optimized constraint set is obtained, and the selected materials are used to generate an optimized house visit video.
[0084] Specifically, firstly, the presentation order of functional areas in the house visit video is extracted, and the similarity with the presentation priority sequence is calculated to obtain a priority parameter. For example, the generated house visit video undergoes quality verification, the presentation order of functional areas is extracted, and the similarity with the presentation priority sequence is calculated to obtain the priority parameter. The shot segmentation and scene tags of the house visit video are analyzed using video analysis tools to identify the actual appearance order of each functional area in the final cut, forming the functional area presentation order. The functional area presentation order is compared with the presentation priority sequence, and the longest common subsequence length is calculated using a sequence matching algorithm. After normalization, a similarity score between 0 and 1 is obtained, which is the priority parameter. For example, if the presentation priority sequence is "living room, master bedroom, kitchen, bathroom," and the actual order is "living room, kitchen, master bedroom, bathroom," then the calculated priority parameter is approximately 0.75.
[0085] Furthermore, based on the presentation order and physical location of the functional areas, a smoothness parameter is obtained. For example, the coordinates of the center point of each functional area's floor plan are extracted from the property feature data. According to the actual presentation order, the Euclidean distance between the center points of each pair of adjacent functional areas is calculated sequentially, resulting in a set of adjacent distance values. Simultaneously, the center point distances between all pairs of functional areas in the property are calculated, and the median of these distances is taken as a preset reasonable distance threshold. The number of adjacent pairs whose adjacent distances are less than this reasonable distance threshold in the actual presentation order is counted, and this number is divided by the total number of adjacent pairs; the resulting ratio is the smoothness parameter. This parameter ranges from 0 to 1; the closer it is to 1, the more compact and reasonable the flow of movement, and the weaker the sense of spatial jump when switching between adjacent functional areas.
[0086] Furthermore, brightness parameters are obtained based on the illumination area sequence of the house visit video. For example, the proportional distribution of highlight and shadow areas in each frame of the house visit video is calculated to generate a sequence curve of illumination area ratio over time. The variance of this sequence is calculated to characterize brightness uniformity. The ratio of this variance to a preset ideal variance is obtained to acquire the brightness parameters. The preset ideal variance can be obtained based on the upper quartile values from historical house visit videos.
[0087] Furthermore, the priority parameter, the smoothness parameter, and the brightness parameter are weighted and calculated to obtain the editing adaptation parameter. For example, three weighting factors are set: priority weight, smoothness weight, and brightness weight, with the sum of each weight being 1. The editing adaptation parameter is calculated as follows: Editing adaptation parameter = Priority parameter × Priority weight + Smoothness parameter × Smoothness weight + Brightness parameter × Brightness weight.
[0088] Furthermore, based on the clip adaptation parameters, the number of optimized sets Q is obtained.
[0089] Specifically, first, the deviation between the obtained editing adaptation parameters and the editing adaptation threshold is calculated. Deviation = |editing adaptation parameters - editing adaptation threshold| / editing adaptation threshold.
[0090] Further, the deviation is multiplied by the preset number of optimization sets to obtain the number of optimization sets Q. This deviation is then multiplied by the preset number of optimization sets to get the number of optimization sets Q, and rounded down to ensure that Q is a positive integer. The Q value represents the number of optimization schemes that need further fine-tuning and filtering. For example, if the preset number of optimization sets is set to 5 and the deviation is 0.3, then Q = 5 × 0.3 = 1.5, and rounded down, the number of optimization sets Q is 1.
[0091] Furthermore, based on the number of optimized sets, multiple material screenings and editing processes are performed to obtain multiple candidate home visit videos, and the candidate home visit video with the largest editing adaptation parameter is selected as the optimized home visit video.
[0092] Specifically, firstly, the editing constraint set and the filtering constraint set are fine-tuned to generate Q optimized editing constraint sets and optimized filtering constraint sets, wherein the fine-tuning range is obtained based on the deviation degree. For example, the fine-tuning range is scaled proportionally according to the deviation degree; the larger the deviation degree, the larger the fine-tuning range. In the editing constraint set, the duration requirements of each functional area are multiplied by a random coefficient, and the priority of transition types is reordered; in the filtering constraint set, the lighting angle tolerance range is scaled, and the minimum duration threshold of the segment is increased or decreased. Each optimized constraint set is randomly sampled from a preset fine-tuning parameter space to ensure that the Q constraint sets are different.
[0093] Further, the Q sets of optimized clipping constraints and the Q sets of optimized filtering constraints are randomly paired to obtain multiple constraint set pairs, wherein each constraint set pair includes one set of optimized clipping constraints and one set of optimized filtering constraints. Specifically, the Q sets of optimized clipping constraints and the Q sets of optimized filtering constraints are randomly paired to obtain Q... 2 Each constraint set pair contains an optimized editing constraint set and an optimized filtering constraint set, which together constitute the instruction set for a complete editing task.
[0094] Furthermore, the constraint set is used to perform multiple material screenings and editing to obtain multiple candidate home visit videos. Specifically, the complete process of the video editing module is repeatedly invoked, and the material library is searched again based on the current constraint set, and the materials are re-edited and spliced to generate a candidate home visit video. This process executes a total of Q... 2 Next, obtain Q 2A total of 10 candidate room visit videos were selected. Further, the aforementioned quality verification process was performed on each candidate video, calculating the editing adaptation parameters for each video. The candidate video with the highest value of the editing adaptation parameters was selected as the final optimized room visit video output. This video achieved an optimal balance in terms of priority matching, smooth movement, and uniform lighting.
[0095] This application embodiment performs multi-dimensional quality verification on house visit videos, enabling the system to possess objective, quantifiable, and interpretable final video quality perception capabilities. Secondly, based on the deviation between the editing adaptation parameters and preset thresholds, the module dynamically calculates the number of optimization sets and uses this as a basis to fine-tune the original screening constraint sets and editing constraint sets, generating multiple optimized versions of constraint rule combinations. Subsequently, the system uses these optimized constraint sets to perform a new round of screening and editing of the materials, generating multiple candidate house visit videos, and selecting the one with the best editing adaptation parameters as the final optimized house visit video output. Through this optimization loop, not only can quality defects exposed during a single generation process be automatically corrected, but rule optimization experience can also be continuously accumulated during continuous operation, resulting in a continuous improvement in the overall video output quality.
[0096] Example 2, as Figure 2 As shown, embodiments of the present invention also provide a method for intelligent generation of real estate exploration videos throughout the entire process based on multimodal AI, including:
[0097] Obtain the property features of the property to be used in the video.
[0098] This includes:
[0099] Obtain the property features of the property to be used in the video, wherein the property features include functional area features, orientation features, and key features.
[0100] Based on the property characteristics, functional area weights are obtained, and lighting conditions are obtained by combining the property characteristics with lighting matching.
[0101] This includes:
[0102] The property features are input into an importance allocation model to obtain functional area importance parameters;
[0103] The configuration of the importance allocation model includes:
[0104] Historical house visit videos were acquired, edited, and analyzed to obtain functional area importance parameters, which were then used as a sample output set.
[0105] Collect the property characteristics of the properties corresponding to the historical house visit videos, and use them as the sample input set;
[0106] Based on machine learning, an importance allocation model is constructed, and the importance allocation model is trained using the sample input set and the sample output set until convergence, thus completing the configuration of the importance allocation model;
[0107] The importance parameters of the multiple functional areas are normalized to obtain the functional area weights;
[0108] By combining the orientation characteristics of the property, the optimal lighting angle can be obtained through analysis.
[0109] Based on the optimal lighting angle and the corresponding relationship of natural lighting time periods, the lighting conditions are derived, wherein the lighting conditions include the shooting time.
[0110] Based on the functional area weights and the lighting conditions, the time sequence of the functional areas is obtained, forming a set of filtering constraints and a set of editing constraints. The selected materials are then edited to generate a house visit video.
[0111] This includes:
[0112] Based on the functional area weights, obtain the presentation priority sequence and duration requirements for each functional area;
[0113] Based on the lighting conditions and duration requirements, obtain the set of filtering constraints;
[0114] Based on the presentation priority sequence and the duration requirement, obtain the editing constraint set;
[0115] The materials are filtered based on the filtering constraint set, and the filtered materials are edited in conjunction with the editing constraint set to generate a house visit video.
[0116] The quality of the house visit video is verified, editing adaptation parameters are obtained, the filtering constraint set and the editing constraint set are optimized, an optimized constraint set is obtained, and the selected materials are used to generate an optimized house visit video.
[0117] This includes:
[0118] Extract the presentation order of the functional areas in the house visit video, calculate the similarity with the presentation priority sequence, and obtain the priority parameter;
[0119] Based on the presentation order and physical location of the functional areas, obtain the fluency parameters;
[0120] Based on the illumination area sequence of the house visit video, brightness parameters are obtained;
[0121] The priority parameter, the smoothness parameter, and the brightness parameter are weighted and calculated to obtain the editing adaptation parameter;
[0122] Based on the editing adaptation parameters, obtain the number of optimized sets Q;
[0123] Among them, obtaining the number Q of optimized sets based on the editing adaptation parameters includes:
[0124] Calculate the deviation between the clip adaptation parameters and the clip adaptation threshold;
[0125] The number of optimization sets Q is obtained by multiplying the deviation by the preset number of optimization sets;
[0126] Based on the number of optimization sets, multiple material screenings and editing processes are performed to obtain multiple candidate house visit videos, and the candidate house visit video with the largest editing adaptation parameter is selected as the optimized house visit video.
[0127] Based on the number of optimized sets, multiple rounds of material screening and editing are performed to obtain multiple candidate house visit videos, including:
[0128] Fine-tuning the editing constraint set and the filtering constraint set generates Q optimized editing constraint sets and optimized filtering constraint sets, wherein the fine-tuning range is obtained based on the deviation degree;
[0129] The Q sets of optimized clipping constraints and optimized filtering constraints are randomly paired to obtain multiple constraint set pairs, wherein each constraint set pair includes an optimized clipping constraint set and an optimized filtering constraint set;
[0130] Multiple material screenings and editing processes were performed using the aforementioned constraint set to obtain multiple candidate house visit videos.
[0131] In summary, the embodiments of this application have at least the following technical effects:
[0132] This application proposes a multimodal AI-based intelligent generation system for the entire process of real estate exploration videos. By deeply integrating multi-dimensional real estate features with intelligent content production processes, it constructs a complete solution for generating real estate exploration videos, encompassing feature analysis, lighting matching, material selection, automated editing, quality verification, and closed-loop optimization. The system first uses a machine learning-trained importance allocation model to intelligently deconstruct the relative importance of different functional areas based on the real estate features to be generated, including functional area features, orientation features, and key features. This model generates normalized functional area weights, providing objective and quantifiable decision-making basis for subsequent content arrangement. Simultaneously, by combining the correspondence between orientation features and natural lighting periods, it automatically matches the optimal lighting angle and shooting time, fundamentally ensuring the professionalism and visual comfort of the video's lighting presentation. Building upon this foundation, the system adaptively generates presentation priority sequences, duration requirements, and filtering and editing constraint sets for functional areas based on functional area weights and lighting conditions. This enables the system to accurately select highly relevant content from source material and automatically edit and arrange it according to professional narrative logic, resulting in highly coordinated room visit videos in terms of functional display completeness, smooth movement, and emphasis on key points. Furthermore, the system performs multi-dimensional quality checks on the initial room visit video, comprehensively calculating editing adaptation parameters from the perspectives of priority matching, spatial movement smoothness, and lighting uniformity. Based on these parameters, the system dynamically optimizes the filtering and editing constraint sets. Through a deviation-driven optimization set generation and constraint set fine-tuning mechanism, the system selects and outputs the highest-quality optimized room visit video from multiple candidate videos. Compared to traditional methods, the technical solution provided in this application significantly improves the automation and efficiency of room visit video production. In addition, the closed-loop quality check and constraint set optimization enables the system to continuously iterate, dynamically adjusting the generation strategy based on the quality feedback of the final video, significantly enhancing the overall quality of the output video.
[0133] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0134] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0135] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A real estate exploration video intelligent generation system based on multimodal AI, characterized in that: include: The feature acquisition module is used to acquire the property features of the property to be generated in the video. The condition acquisition module is used to obtain the functional area weights based on the property features, and to perform lighting matching in combination with the property features to obtain the lighting conditions. The video editing module is used to obtain the time sequence of the functional areas based on the functional area weights and the lighting conditions, form a filtering constraint set and an editing constraint set, and filter the materials to edit and generate a house visit video. The editing optimization module is used to perform quality verification on the house visit video, obtain editing adaptation parameters, optimize the filtering constraint set and the editing constraint set, obtain the optimized constraint set, and filter materials to generate an optimized house visit video.
2. The intelligent generation system for the entire process of real estate exploration video based on multimodal AI as described in claim 1, characterized in that, Obtain the property features of the property for which the video to be generated, including: Obtain the property features of the property to be used in the video, wherein the property features include functional area features, orientation features, and key features.
3. The intelligent generation system for the entire process of real estate exploration video based on multimodal AI as described in claim 1, characterized in that, Based on the property characteristics, functional area weights are obtained, and lighting conditions are obtained by combining the property characteristics with lighting matching, including: The property features are input into an importance allocation model to obtain functional area importance parameters; The importance parameters of the multiple functional areas are normalized to obtain the functional area weights; By combining the orientation characteristics of the property, the optimal lighting angle can be obtained through analysis. Based on the optimal lighting angle and the corresponding relationship of natural lighting time periods, the lighting conditions are derived, wherein the lighting conditions include the shooting time.
4. The intelligent generation system for the entire process of real estate exploration video based on multimodal AI as described in claim 3, characterized in that, The configuration of the importance allocation model includes: Historical house visit videos were acquired, edited, and analyzed to obtain functional area importance parameters, which were then used as a sample output set. Collect the property characteristics of the properties corresponding to the historical house visit videos, and use them as the sample input set; Based on machine learning, an importance allocation model is constructed, and the importance allocation model is trained using the sample input set and the sample output set until convergence, thus completing the configuration of the importance allocation model.
5. The intelligent generation system for the entire process of real estate exploration video based on multimodal AI according to claim 1, characterized in that, Based on the functional area weights and the lighting conditions, a time series sequence of the functional areas is obtained, forming a filtering constraint set and an editing constraint set. The selected materials are then edited to generate a house visit video, including: Based on the functional area weights, obtain the presentation priority sequence and duration requirements for each functional area; Based on the lighting conditions and duration requirements, obtain the set of filtering constraints; Based on the presentation priority sequence and the duration requirement, obtain the editing constraint set; The materials are filtered based on the filtering constraint set, and the filtered materials are edited in conjunction with the editing constraint set to generate a house visit video.
6. The intelligent generation system for the entire process of real estate exploration video based on multimodal AI according to claim 1, characterized in that, The quality of the house visit video is verified, and editing adaptation parameters are obtained, including: Extract the presentation order of the functional areas in the house visit video, calculate the similarity with the presentation priority sequence, and obtain the priority parameter; Based on the presentation order and physical location of the functional areas, obtain the fluency parameters; Based on the illumination area sequence of the house visit video, brightness parameters are obtained; The priority parameter, the smoothness parameter, and the brightness parameter are weighted and calculated to obtain the editing adaptation parameter.
7. The intelligent generation system for the entire process of real estate exploration video based on multimodal AI as described in claim 1, characterized in that, Optimize the filtering constraint set and the editing constraint set to obtain an optimized constraint set, and filter the materials to generate an optimized house visit video, including: Based on the editing adaptation parameters, obtain the number of optimized sets Q; Based on the number of optimization sets, multiple material screenings and editing processes are performed to obtain multiple candidate house visit videos, and the candidate house visit video with the largest editing adaptation parameter is selected as the optimized house visit video.
8. The intelligent generation system for the entire process of real estate exploration video based on multimodal AI according to claim 7, characterized in that, Based on the editing adaptation parameters, the number of optimized sets Q is obtained, including: Calculate the deviation between the clip adaptation parameters and the clip adaptation threshold; The number of optimization sets Q is obtained by multiplying the deviation by the preset number of optimization sets.
9. The intelligent generation system for the entire process of real estate exploration video based on multimodal AI according to claim 8, characterized in that, Based on the number of optimized sets, multiple rounds of material screening and editing are performed to obtain multiple candidate house visit videos, including: Fine-tuning the editing constraint set and the filtering constraint set generates Q optimized editing constraint sets and optimized filtering constraint sets, wherein the fine-tuning range is obtained based on the deviation degree; The Q sets of optimized clipping constraints and optimized filtering constraints are randomly paired to obtain multiple constraint set pairs, wherein each constraint set pair includes an optimized clipping constraint set and an optimized filtering constraint set; Multiple material screenings and editing processes were performed using the aforementioned constraint set to obtain multiple candidate house visit videos.
10. A method for intelligent generation of real estate exploration videos throughout the entire process based on multimodal AI, characterized in that: include: Obtain the property features of the property to be used in the video; Based on the property characteristics, functional area weights are obtained, and lighting conditions are obtained by combining the property characteristics with lighting matching. Based on the functional area weights and the lighting conditions, the time sequence of the functional areas is obtained, forming a set of filtering constraints and a set of editing constraints. The selected materials are then edited to generate a house visit video. The quality of the house visit video is verified, editing adaptation parameters are obtained, the filtering constraint set and the editing constraint set are optimized, an optimized constraint set is obtained, and the selected materials are used to generate an optimized house visit video.