House resource video processing method and device, and storage medium
By using artificial intelligence technology to intelligently process property listing videos throughout the entire process, the problems of low efficiency in property listing video generation and dull user experience have been solved. This has enabled the automated generation and publishing of high-quality property listing videos, improving user experience and update efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-10
AI Technical Summary
The existing property listing videos are generated inefficiently, resulting in a dull user experience. Post-processing by property providers is time-consuming and demanding, and updates are untimely and cumbersome.
By leveraging artificial intelligence technology for intelligent processing of property listing videos throughout the entire process, including spatial recognition, visual understanding, large language model generation of explanatory text, and audio synthesis, the system enables automated generation and publishing of high-quality target property listing videos from initial videos.
Significantly improves the efficiency and quality of property listing video generation, provides a better user experience, supports automatic video updates, and reduces the shooting difficulty and post-processing burden for property listing providers.
Smart Images

Figure CN121640343A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a house source video processing method, device and storage medium. BACKGROUND
[0002] In the field of real estate services, when a user has a demand for renting or buying a house, the user can browse a large number of house sources through a platform. A traditional house source display mode mainly displays static pictures. However, the display mode of static pictures has a problem of information fragmentation, and cannot meet the immersive experience demand of the user. Therefore, the platform provides a video house viewing function, and the user can view house source videos online. The videos include videos of each room in the house source, so that the user can more intuitively and truly view and understand the situation of the house. The video house viewing function not only can meet the online house viewing demand of the user, but also can save the time cost of offline house viewing of the user.
[0003] However, at present, the original house source video is shot by a house source provider. The pure house source video makes the user feel boring, and post-processing of the original video needs the house source provider to spend a large amount of time. The video generation efficiency is low, and the post-processing ability requirement of the house source provider is high. SUMMARY
[0004] The embodiments of the present application provide a house source video processing method, device and storage medium, which can realize the whole process publishing of the house source video based on artificial intelligence, improve the house source video generation efficiency, and improve the user experience.
[0005] In a first aspect, embodiments of this application provide a method for processing video of a property, comprising: acquiring an initial video of a target property captured on camera, the initial video comprising multiple video frames, and the target property comprising multiple spatial regions; utilizing a spatial recognition model, based on the temporal correlation between video frames and structured prior knowledge of the indoor scene, performing object detection and spatial category reasoning on the video frames to obtain preliminary spatial segmentation results, and performing secondary fine-grained correction on the boundaries of the spatial regions based on the preliminary spatial segmentation results to obtain spatial recognition results, the spatial recognition results including the start and end times of the video frames corresponding to the recognition region of each space, the spatial category label, and the objects contained in the space; utilizing a large-scale visual understanding model, from multiple Frames that meet the requirements are selected from the video frames as cover frames. Using the first language model, items contained in each space are used as material, and the start and end times of the video frames corresponding to the recognition areas of each space are used to limit the length of the explanatory text for each space. Through a multi-round structured reasoning mechanism, the intermediate results of the previous reasoning output are used to constrain and optimize the text, generating explanatory text that is aligned with the content of the initial video. The explanatory text is input into the speech synthesis engine to obtain the corresponding audio, which is used to play synchronously with the initial video. The initial video, cover frame, and audio are synthesized to generate the target property video, which is then published to the target application for users to watch.
[0006] Secondly, embodiments of this application also provide an electronic device, including: a memory and a processor; wherein, the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the method described in the first aspect.
[0007] Thirdly, embodiments of this application provide a housing video processing apparatus, including: The acquisition module is used to acquire the initial video of the target property, which includes multiple video frames and the target property includes multiple spatial areas. The spatial recognition module is used to perform object detection and spatial category reasoning on video frames based on the temporal correlation between video frames and the structured prior knowledge of indoor scenes using a spatial recognition model. This results in preliminary spatial segmentation results, and secondary fine-grained correction of the boundaries of spatial regions is performed based on the preliminary spatial segmentation results to obtain spatial recognition results. The spatial recognition results include the start and end times of the video frames corresponding to each spatial recognition region, the spatial category label, and the objects contained in the space. The corpus generation module uses a large visual understanding model to select a frame that meets the requirements from multiple video frames as the cover frame. Using the first large language model, the module uses the items contained in each space as material and limits the length of the explanatory text for each space to the start and end times of the video frames corresponding to the recognition areas of each space. Through a multi-round structured reasoning mechanism, the module optimizes the constraints based on the intermediate results of the previous reasoning output round by round to generate explanatory text that is aligned with the content of the initial video. The audio generation module is used to input the explanatory text into the speech synthesis engine to obtain the corresponding explanatory audio, which is used to play synchronously with the initial video footage. The synthesis and publishing module is used to synthesize the initial video, cover frame, and explanatory audio to generate the target property video, and then publish the target property video to the target application for users to watch.
[0008] Fourthly, embodiments of this application provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the method described in the first aspect.
[0009] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program or instructions, which, when executed by a processor, cause the processor to implement the method described in the first aspect above.
[0010] The method provided in this application acquires an initial video of the target property and inputs it into a spatial recognition model. This model, based on the temporal correlation between video frames and structured prior knowledge of the indoor scene, first performs object detection and spatial category reasoning to obtain preliminary spatial segmentation results. Then, it performs secondary fine-grained correction on the spatial region boundaries, thereby accurately segmenting the space of the initial video. On this basis, a large-scale visual understanding model intelligently selects cover frames that meet visual quality requirements from the video frames. The spatial recognition results are then input into a first large-scale language model, which uses objects in each space as generation materials and strictly limits the duration of the corresponding explanatory text based on the start and end times of the video frames corresponding to the recognition areas of each space. Through a multi-round iterative reasoning mechanism, it generates explanatory text whose semantic content, narrative rhythm, and time span are precisely aligned with the initial video footage. Subsequently, the explanatory text is converted into audio by a speech synthesis engine and synthesized with the initial video and cover frame to generate the target property video, which is then published to the target application. This achieves intelligent processing throughout the entire process, from raw video processing and multi-dimensional content generation to automatic video publishing, significantly improving the efficiency and quality of property video generation.
[0011] This technical solution achieves fully automated, high-fidelity conversion from initial video to high-quality target property video through a closed-loop process involving spatial recognition, structured temporal information extraction, multi-round constraint-driven text generation, and synchronized audio-visual synthesis. Specifically, the spatial recognition model significantly improves the temporal and semantic accuracy of spatial segmentation by jointly utilizing temporal correlation and prior knowledge, along with secondary boundary correction. Meanwhile, the large language model, under the hard constraint of video frame start and end times, combined with a constraint transmission mechanism based on multi-round structured reasoning, effectively avoids common problems in traditional text generation, such as explanation timeouts, content gaps, or omissions of key items. This ensures that the explanation text not only semantically matches the actual layout and furnishings of each space, but also that its audio playback duration matches the time period of the corresponding space in the video, providing a foundation for audio-visual synchronization. Finally, the synthesized property video achieves a high degree of synchronization and consistency between scene transitions, audio explanations, and spatial semantics, significantly improving the user's information acquisition efficiency and immersive experience, resulting in a better viewing experience. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a method for processing property video provided in an embodiment of this application; Figure 2 A flowchart illustrating another method for processing property video provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0015] In the real estate services sector, users with rental or purchase needs can browse a vast number of listings on the platform. Traditionally, listings are displayed primarily through static images. However, static images suffer from fragmented information and fail to meet users' needs for an immersive experience. To address this, the platform offers a video viewing feature, allowing users to watch videos of individual rooms within the property. This provides a more intuitive and realistic view of the property, satisfying users' online viewing needs while saving them the time and effort of in-person visits.
[0016] However, currently, the original videos for property listings are shot by the property providers themselves. These simple videos can be monotonous for users, and post-processing the original videos (such as adding audio, cover art, subtitles, and background music) requires a significant investment of time from the property providers, resulting in low video generation efficiency and demanding high post-processing skills. Furthermore, after a property is published, updating the video content requires the property provider to reprocess it, leading to untimely and cumbersome updates.
[0017] In view of this, this application provides a method for processing property listing videos. This method utilizes artificial intelligence (AI) technology to perform in-depth secondary processing on the original video, automatically generating one or more post-production content elements such as audio narration, subtitles, covers, background music, endings, and supplementary explanations. This results in a high-quality target property listing video containing diverse content, and the video is automatically published. This achieves fully intelligent processing from original video processing and multi-dimensional content generation to automatic video publishing, significantly improving the efficiency and quality of property listing video generation. The method provided in this application allows the property listing photographer to focus on shooting the original video without needing to provide narration while shooting, reducing the difficulty of shooting the original video and improving video quality. Furthermore, by generating clear and content-rich audio narration, users can quickly and comprehensively understand the property information; the addition of subtitles, covers, endings, background music, and supplementary explanations to the target property listing video enhances the user's viewing experience.
[0018] Furthermore, after the target property video is published, this application embodiment also provides a video update module, which can dynamically update the explanation content and cover of the target property video based on user questions and answers, thereby realizing automatic updating of the property video.
[0019] The method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0020] Figure 1 A flowchart illustrating a method for processing property video according to an embodiment of this application is shown. Figure 1 As shown, the method may include steps S101-S106.
[0021] S101, acquire the initial video of the target property, which includes multiple video frames and the target property includes multiple spatial areas.
[0022] S102. Using a spatial recognition model, based on the temporal correlation between video frames and the structured prior knowledge of indoor scenes, object detection and spatial category reasoning are performed on the video frames to obtain preliminary spatial segmentation results. Based on the preliminary spatial segmentation results, secondary fine-grained correction is performed on the boundaries of the spatial regions to obtain spatial recognition results. The spatial recognition results include the start and end times of the video frames corresponding to the recognition region of each space, the spatial category label, and the objects contained in the space.
[0023] S103 uses a large visual understanding model to select a frame that meets the requirements from multiple video frames as the cover frame.
[0024] S104 utilizes the first major language model, using the items contained in each space as material, and limits the length of the explanatory text for each space to the start and end times of the video frames corresponding to the recognition areas of each space. Through a multi-round structured reasoning mechanism, it performs constraint optimization based on the intermediate results of the previous reasoning output round by round, generating explanatory text that is aligned with the content of the initial video.
[0025] S105 inputs the explanatory text into the speech synthesis engine to obtain the corresponding explanatory audio, which is used to play synchronously with the initial video footage.
[0026] S106, synthesize the initial video, cover frame and explanatory audio to generate the target property video, and publish the target property video to the target application for users to watch.
[0027] In this embodiment, the property provider can provide the initial video footage captured of the target property to a computing device as video material for subsequent production of a video of the target property. There are no restrictions on the content provider; for example, it could be a real estate agent or landlord, or other relevant personnel. Furthermore, there are no limitations on the type of computing device; it can be a desktop computer, laptop computer, tablet computer, or smartphone, or a server-side device such as a conventional server, cloud server, or server array.
[0028] For example, such as Figure 2 As shown, the agent films the target property to obtain an initial video and uploads it to the target application on the first client. Within the target application, the agent selects to create a video of the target property. Then, the target application sends a video creation request, which includes the initial video, to the computing device on the server. Upon receiving the video creation request, the computing device uses AI algorithms to process the initial video, including generating explanatory text, audio, subtitles, background music, a cover image, an ending, property information descriptions, and explanations of the surrounding neighborhood, thus obtaining the target property video. The initial video can be a single-shot video or a multi-shot video.
[0029] Upon receiving a video production request, the computing device first uses a spatial recognition model to perform spatial recognition and segmentation on the initial video, identifying multiple spaces included in the initial video, such as the living room, bedroom, kitchen, bathroom, and balcony. These spatial regions can also be referred to as spatial units, spatial segments, or spaces. The spatial recognition process is explained in detail below.
[0030] The initial video is input into the spatial recognition model. The preprocessing module of the spatial recognition model performs frame extraction preprocessing on the initial video to obtain a target frame sequence composed of multiple high-quality video frames. For example, the initial video stream can be sliced second by second, dividing the initial video into time windows in seconds. Within each time window, multiple target frames that meet the requirements are selected from the original frame sequence (e.g., 30 frames / second). The multiple target frames extracted from multiple time windows form the target frame sequence.
[0031] For example, multiple target frames that meet the requirements can be filtered based on image integrity, sharpness, and brightness. Image integrity is used to assess whether the main scene elements in the image are occluded or segmented, and can be measured by the proportion of the occluded / segmented area to the total area of the image in the video frame. If the proportion of the occluded / segmented area to the total area of the image in a video frame is greater than or equal to a first threshold (e.g., 30%), then the image integrity of that video frame is poor. Sharpness represents the sharpness and detail richness of a video frame, and can be measured by the second derivative (Laplacian) variance. If the Laplacian variance of a video frame is less than a second threshold (adaptively adjusted according to the overall brightness of the video), then the sharpness of that video frame is poor. Brightness represents the overall brightness of the image, and can be evaluated by the mean value of the Y channel in the YUV color space. If a video frame is overexposed (mean Y > 230) or underexposed (mean Y < 40), then the brightness of that video frame is poor. Overexposure refers to a mean Y channel value greater than the threshold (e.g., > 230), and underexposedness refers to a mean Y channel value less than the threshold (e.g., < 40). Within each time window, a picture integrity score, a sharpness score, and a brightness score are calculated for each video frame. The quality score of each video frame is obtained by weighting these three dimensions. The top N frames with the highest quality scores are selected as multiple target frames that meet the requirements, where N ≥ 1. For example, N can be 2, indicating that 2 frames are extracted per second.
[0032] By extracting multiple high-quality video frames in each time window, a sequence of target frames arranged in chronological order is obtained and used as input data for the spatial recognition model, thereby reducing redundant computation while ensuring temporal resolution.
[0033] The target frame sequence is input into the analysis engine of the spatial recognition model. This engine performs fine-grained indoor object analysis on each target frame, accurately identifying various types of furniture and appliances, such as sofas, coffee tables, televisions, TV cabinets, beds, air conditioners, refrigerators, and kitchenware, and outputs the location and confidence level of each type of furniture and appliance. Based on the detected items, the model combines a pre-defined item-space mapping knowledge base. The spatial recognition model automatically weights the discrimination contribution of different items through an attention mechanism, integrates item semantics and global scene context, and infers the spatial category of each frame. It outputs the probability distribution of the indoor spatial category for each frame, obtaining the spatial prediction sequence. To enhance temporal coherence, the model introduces temporal position encoding and cross-frame attention mechanisms to implicitly model the dynamic association between adjacent frames, enabling single-frame prediction to have context-aware capabilities and providing a foundation for subsequent temporal consistency analysis. Among them, the item-space mapping knowledge base encodes strong association rules between typical item combinations and spatial categories. For example, the combination of a stove, range hood, and cabinet represents a kitchen, and the combination of a bed and wardrobe represents a bedroom.
[0034] Furthermore, the spatial recognition model employs a sliding window mechanism to perform batch analysis on the spatial prediction sequences output by the analysis engine. Specifically, a window length is set, and the window slides along the time axis with a certain step size, allowing multiple frames to overlap between adjacent windows. During the sliding process, the model checks whether the spatial categories of all frames within each window are consistent: if the spatial categories of all frames are completely consistent, the window is determined to be in a stable spatial segment; if there is a category change, the frame position where the category first changes is marked as a potential spatial switching point, and the second-level time of the switching point is recorded. For example, in a certain window, the first 6 frames are all living room, but from the 7th frame onwards it changes to kitchen, then the system marks the 7th frame as a potential spatial switching point and records the second-level time of the 7th frame as the 5.0 second of the initial video.
[0035] For identified stable spatial segments (i.e., those with no category differences within the window), a spatial topology confidence assessment is further performed. For example, an interior spatial layout knowledge graph is pre-constructed, encoding the physical connection patterns between functional areas in common residences (e.g., the master bedroom is usually larger than the secondary bedroom, the TV cabinet is adjacent to a wall, etc.). The spatial category of the current stable segment is matched with the topological rules in the interior spatial layout knowledge graph, and its topological rationality score is calculated. If the score is lower than a preset threshold, the spatial category confidence of that segment is reduced, and it is marked as an area requiring review; otherwise, the original prediction result is retained as a high-confidence output.
[0036] For regions near the marked potential spatial switching points (i.e., windows where class differences exist), multi-model collaborative validation is further performed. For example, an integrated system consisting of multiple heterogeneous spatial recognition models is pre-built, including one main backbone model and nine lightweight auxiliary models with different structures (such as spatial classifiers finely tuned based on architectures such as EfficientNet-B3, ResNet-50, and ViT-Tiny). When the integrated system verifies potential spatial switching points, it expands a time window by 1 second before and after the switching point and selects high-density frames from it. Then, all models make independent predictions on these high-density frames, and each model outputs the spatial category it believes is most likely. The number of model votes for each spatial category in each frame is counted, i.e., whether multiple models vote for the same spatial category in the same frame. If a spatial category (e.g., living room) in a frame receives ≥80% (i.e., at least 8 models) of consistent votes, then the category is determined to be the true spatial category of that frame. Based on this, the system traces back to the frame position where the category (e.g., living room) first becomes the "dominant category" in the time series. That is, starting from the beginning frame of the window, the system checks frame by frame to find the first starting frame that satisfies "the category (e.g., living room) continues to receive ≥80% model support in multiple consecutive frames thereafter". The time point corresponding to this frame is determined as the spatial switching moment, and candidate spatial switching points are obtained, resulting in preliminary spatial segmentation results. Otherwise, the region is marked as a low-confidence conflict zone and enters the manual review queue or triggers a rollback strategy.
[0037] All candidate switching points are input into the precision enhancement decision-maker of the spatial recognition model. In the time intervals adjacent to the switching points identified in the initial spatial segmentation results, video frames are supplemented with a higher temporal granularity than the conventional sampling frequency to perform sub-second segmentation. For each candidate switching point, the precision enhancement decision-maker intensively extracts multiple frames at 0.1-second intervals within a 1-second time interval centered on the candidate switching point. These frames are then input again into the analysis engine of the spatial recognition model for spatial recognition, yielding fine-grained spatial probability outputs. Subsequently, the precision enhancement decision-maker employs a first-order difference maximum detection algorithm to calculate the abrupt change in spatial category probability between adjacent 0.1-second frames, determining the moment of the most dramatic change as the final spatial segmentation boundary. Thus, the spatial recognition model can output accurate spatial recognition and segmentation results in 0.1-second time units. The spatial recognition results include the start and end timestamps of the video frames corresponding to each spatial segment, the spatial category label, and the items contained in each space. For example, the model outputs {a list of spaces: [living room, master bedroom, kitchen...], items in each space: {living room: [L-shaped sofa, TV cabinet, bay window tatami], time periods for each space}.
[0038] Optionally, during the aforementioned video spatial recognition process, this embodiment can also automatically identify spatial highlights to label high-value attribute features of each spatial segment in real time during spatial recognition. For example, the spatial recognition model includes a lightweight highlight attribute detection subnetwork, which is deeply coupled with the analysis engine of the spatial recognition model. The highlight recognition process shares underlying features with object recognition and spatial recognition, synchronously completing the detection and association of highlight features at the frame-level semantic understanding stage, achieving end-to-end joint recognition of spatial categories and highlight attributes. Specifically, during the process of the analysis engine identifying the spatial category of the target frame, the highlight attribute detection subnetwork is activated. This subnetwork performs multi-task feature inference strongly correlated with spatial semantics based on the multi-scale feature map output by the analysis engine, outputting the highlight attributes possessed by the space. For example, when the analysis engine identifies a frame as belonging to a bathroom, the highlight attribute detection subnetwork can determine whether the bathroom has windows. If the result is "lit bathroom," it can be labeled "[lit bathroom]" in the spatial category attribute of that frame. Similarly, when identifying frames of a living room or bedroom, the highlight attribute detection subnetwork can infer the presence of specific architectural or decorative features such as floor-to-ceiling windows based on overall visual characteristics, and label them "[floor-to-ceiling window]" in the spatial category attribute. Through highlight analysis, accurate video segmentation can be achieved while automatically mining and structured outputting spatial highlights, providing a material foundation for narration corpora.
[0039] Optionally, after spatial segmentation is completed, structured reasoning across spatial segments can be performed based on all spatial segments and their attributes to generate high-level information reflecting the physical structure and configuration level of the entire housing unit.
[0040] In this embodiment, based on the spatial segmentation results of the initial video, the overall selling points of the property can be further mined, achieving a leap from "knowing which room each frame belongs to" to "understanding what the entire house looks like, how it is laid out, and how it is configured." For example, the inference engine of the spatial recognition model can perform cross-segment association, geometric estimation, and floor plan deduction based on the spatial category of each spatial segment, the start and end times of the corresponding video frames, and the items contained therein. This completes high-level semantic understanding tasks such as intelligent spatial determination, room opening recognition, floor plan alignment, and furniture extraction, outputting structured property selling point information including floor plan diagrams, number and type of rooms, room opening layout, and furniture configuration level. For example, the structured property information could be: {Floor plan: three bedrooms, two living rooms, two bathrooms; North-south facing: true; Living room width: 3.9m; Furniture: fully furnished}.
[0041] Among these features, the intelligent spatial identification refers to recognizing whether multiple recurring or visually coherent spatial segments belong to the same physical room. Since a single-shot video may traverse the same physical space multiple times (e.g., shooting from the left side of the living room to the right, then circling back to the sofa area), the model cannot simply treat all frames labeled "living room" as continuous segments. Therefore, the inference engine can perform spatial aggregation based on spatial geometric consistency and viewpoint overlap: for all segments identified as belonging to the same category (e.g., living room), its corresponding visual features (e.g., wall texture, main furniture layout, window position) and temporal trajectory (frame sequence continuity, motion direction) are extracted. By calculating the feature similarity and spatial connectivity between segments, it determines whether these segments belong to the same physical room. If determined to be the same space, they are merged into a single logical spatial unit, and their total duration, covered viewpoint range, and representative images are recorded. This avoids misjudgments such as "a living room being split into three segments" due to complex shooting paths, providing accurate spatial counting for subsequent floor plan reconstruction.
[0042] Wide-span recognition refers to estimating the width of the light-receiving surface of major functional areas (such as the living room and master bedroom) based on the temporal adjacency between spatial segments, object co-occurrence patterns, and geometric priors, in order to quantify the openness of the space. After identifying major functional areas such as the living room and master bedroom, the floor plan inference engine automatically performs wide-span estimation: First, it uses monocular depth estimation to recover the 3D structure of the scene from representative frames; second, it combines the identified window or balcony boundaries to extract the maximum unobstructed span of the space in the horizontal direction; and then, through the in-camera participation perspective geometry model, it converts the pixel span into the actual physical width. For example, if a continuous floor-to-ceiling window is detected in the living room video segment, covering 70% of the screen width, and the depth map shows that the wall is about 4 meters away from the foreground sofa, then the estimated wide-span value is 3.8 ± 0.3 meters. Based on this detection result, property information for a "large open-plan living room" can be generated.
[0043] Apartment layout alignment refers to inferring the topological connections between rooms and the overall apartment layout by utilizing the connection order between various spatial segments. Based on all spatial units and their interconnections, the apartment layout inference engine can construct a dynamic apartment layout sketch: each spatial unit serves as a graph node, with node attributes including type, area estimation, room width, and whether it has a window in the bathroom, etc.; the edges between nodes are generated by the actual spatial transition order traversed in the video, and are combined with a spatial topology knowledge graph for rationality correction; subsequently, the apartment layout inference engine performs graph matching between this sketch and a pre-stored library of common standardized apartment layout templates, calculates structural similarity, and outputs the name of the most matching standard apartment layout, thus automatically inferring the existence of areas not explicitly filmed.
[0044] Furniture matching extraction refers to aggregating the detection results of items in each spatial segment, analyzing the completeness and quality of furniture configuration in functional areas, and forming a spatial functional matching profile. Within each spatial segment, the output of the item detection head of the analysis engine of the spatial recognition model is reused to not only identify major categories such as beds and sofas, but also further determine the furniture style (such as modern minimalist, new Chinese style), matching completeness, and functional configuration (such as whether the kitchen includes a built-in dishwasher). This information is structurally aggregated through a predefined furniture matching rule base: for example, if a double bed, double bedside tables, a walk-in wardrobe, and a dressing table are detected simultaneously in the master bedroom, it is marked as a fully furnished master bedroom.
[0045] Through the above spatial segmentation, brightness extraction, and property information mining, the final spatial recognition model can output the start and end timestamps of the video frames corresponding to the recognition area of each space, the space category label, the items contained in each space, the space brightness attributes, and the property's selling points, providing a material basis for the narration corpus.
[0046] The training process, architecture, and processing flow of the spatial recognition model are illustrated below.
[0047] In this embodiment of the application, the spatial recognition model uses approximately 10 9 (1B) Based on a backbone network of parameter magnitude, the system achieved dynamic recognition of indoor spatial structures and sub-second semantic segmentation capabilities through end-to-end fine-tuning training on a dataset of over 10,000 real housing video sets. The housing video dataset covers videos of different apartment types, decoration styles, and lighting conditions in major cities across the country. Each video contains synchronous spatiotemporal annotations, including: (1) spatial category labels for each frame (living room, master bedroom, kitchen, bathroom, balcony, entrance hall, dining room, etc.); (2) precise timestamps of spatial switching points (accuracy 0.1 seconds); (3) bounding boxes and categories of key items; (4) spatial brightness attribute labels and housing selling point information labels (such as bright bathroom, north-south facing); (5) apartment structure topology. This annotation system provides a supervisory signal for the spatial recognition model to learn spatial segmentation, item recognition, highlight recognition, and housing information recognition simultaneously. By introducing spatial highlight attributes as supervisory signals during the model training stage, the lightweight highlight attribute detection subnetwork automatically establishes the mapping capability from visual features to high-order semantic attributes during the learning process.
[0048] The analysis engine of a spatial recognition model may include a shared visual encoder, a spatiotemporal location embedding module, a temporal context fusion module, and a multi-task decoder.
[0049] The shared visual encoder can be implemented based on the Vision Transformer Base (ViT-Base) architecture and can contain multiple Transformer encoding layers. After the target frame is input into the analysis engine, it is divided into non-overlapping image patches (e.g., 16×16) by the shared visual encoder. Each image patch is linearly projected into a multi-dimensional (e.g., 768) embedding vector. Subsequently, a learnable classification [CLS] label embedding is concatenated before the sequence of multiple image patch embeddings to form the initial visual input sequence. After the initial visual input sequence is processed by the multi-layer encoder, the output includes global semantic features at the [CLS] positions and local features at each image patch position. The global semantic features are used for spatial categories, and the local features are used for item detection.
[0050] The spatiotemporal location embedding module superimposes a learnable temporal location code onto the [CLS] feature of each frame. This code is generated by linear projection of the absolute timestamp in the video. At the same time, it injects the two-dimensional spatial coordinates of each image patch embedding as a positional bias to enhance the model's perception of the spatial layout of objects. The frame-level feature sequences of multiple consecutive frames (corresponding to sliding windows) after spatiotemporal embedding are concatenated into a long sequence and input into the temporal context fusion module.
[0051] The temporal context fusion module consists of a multi-layer Transformer encoder. It uses a self-attention mechanism to globally model the entire long sequence, outputting a context-enhanced feature sequence. Each frame corresponds to an enhanced [CLS] global feature (representing the overall spatial semantics) and multiple enhanced local image patch features (preserving spatial distribution information of objects), enabling the model to capture transitional relationships between spaces. For example, if a stove is detected in frame i and a dining table is detected in frame j, the attention weights will enhance the semantic association between the two on the kitchen-dining room transition path.
[0052] The multi-task decoding head comprises two parallel branches: a spatial classification head and an object detection head. The spatial classification head can employ a two-layer perceptron (MLP) structure, receiving the enhanced [CLS] global image and dynamically fusing object detection results with predefined object-space mapping rules during inference, outputting spatial category probabilities. The object detection head consists of a multi-layer Transformer encoder, receiving enhanced local image patch features and outputting bounding boxes and confidence scores for indoor objects.
[0053] The spatial recognition model's inference engine is based on a unified vision-temporal encoder at its core, which can fuse RGB images of video frames, monocular depth estimation, and timestamp information to output multi-scale spatial semantic features with temporal alignment. The middle layer contains a task collaborative inference module, which uses graph neural networks to model the start and end times of each spatial segment (node) and its corresponding video frames, as well as the co-occurrence relationships (edges) of items, to achieve cross-segment association and intelligent judgment within the same space. The upper layer integrates multiple dedicated sub-headers, including a room opening estimation head, a unit type alignment head, and a furniture matching analysis head, to jointly complete the end-to-end mapping from segmentation results to structured housing selling points.
[0054] In the aforementioned spatial recognition and segmentation process, since the video consists of a large number of temporal frames, any deviation in any frame will lead to inaccurate overall segmentation. This application's embodiments construct a temporally consistent inference framework based on a sliding window. By jointly modeling multiple frames within the sliding window and introducing a cross-frame self-attention mechanism, the model can tolerate misjudgments caused by occlusion or jitter in individual frames. Switching point marking is only triggered when the category change forms a stable trend within the window. Furthermore, the multi-model collaborative verification mechanism further filters out accidental errors of single models in edge scenes, ensuring that local frame-level errors do not directly propagate to the global segmentation result, thereby significantly improving the robustness and accuracy of the overall segmentation.
[0055] To achieve precise segmentation at the millisecond level while simultaneously satisfying the dual constraints of temporal and spatial consistency, this application innovatively introduces a two-stage localization architecture: the first stage achieves second-level initial screening through preprocessing and frame extraction; the second stage only initiates sub-second-level intensive resampling and boundary sharpening algorithms near candidate switching points. This ensures overall efficiency while precisely focusing computational resources on key transition regions. During the fine-tuning stage, the spatial recognition model's analysis engine is injected with temporal location encoding and cross-frame attention, enabling it to implicitly learn the temporal patterns of smooth spatial transitions and concentrated category changes during frame-level prediction. This ensures that even at a 0.1-second granularity, the spatial labels of adjacent segments maintain semantic coherence, avoiding high-frequency jitter.
[0056] After completing spatial segmentation, brightness extraction, and property information mining, the spatial category labels output by the spatial recognition model, the items contained in each space, the spatial brightness attributes, and the property's selling points can be used as source material. The length of the explanatory text for each space is limited by the start and end timestamps of the video frames corresponding to each spatial segment, generating explanatory text highly aligned with the video content. Optionally, the timbre of the narration can also be selected, with different timbres corresponding to different speaking speeds and rhythms. The speaking speed and rhythm can be combined with the duration of each spatial segment to limit the length of the explanatory text.
[0057] In one exemplary embodiment, a multi-round iterative strategy can be employed to generate the explanatory text. This strategy integrates temporal constraints and multi-dimensional content quality criteria during the generation process to ensure that the explanatory text is highly aligned with the initial video content in terms of both timing and semantics. Specifically, the first large language model (LLM) executes a first round of comprehensive generation reasoning and multiple rounds of specialized evaluation and correction reasoning sequentially based on preset round-specific prompts, forming an iterative optimization chain. The first round integrates property information, plans the storyboard content, plans the speaking speed, and generates the explanatory content. Subsequent rounds respectively detect the explanatory duration, naturalness and coherence, visual relevance, content compliance and legality, and information accuracy. Each round of reasoning outputs a structured intermediate result in natural language form, which serves as the input constraint for the next round of reasoning. The final output is an explanatory script that meets the requirements of duration matching, natural speech, synchronized visuals, content compliance, and information accuracy.
[0058] Taking a three-round reasoning approach as an example, in the first round, the spatial recognition and segmentation results are input into a pre-constructed first prompt word template to generate the first prompt word. Guided by the first prompt word, the first large language model plans the focus of explanation for each space, calculates the maximum number of characters allowed for each spatial segment based on the speech rate and rhythm, and forcibly inserts semantic pause markers [PAUSE] within 0.3 seconds before and after the scene switching boundary. The first version of the explanation text is output, which is a sequence of paragraphs with timestamps. The first line of each paragraph indicates the start time point of the text content in the video. Each space corresponds to a paragraph, and the corresponding semantic content references the highlight attributes or items of that spatial segment, and the speech synthesis duration of the paragraph does not exceed the video duration of that spatial segment.
[0059] The second round is a multi-dimensional quality collaborative optimization inference round, used to detect the narration duration, naturalness and coherence, relevance to the visuals, and preliminary compliance screening. In the second round, the first version of the narration text and the pre-synthesized duration report are input into a pre-built second prompt word template to generate second prompt words. Guided by the second prompt words, the first large language model performs multiple collaborative corrections on the first version of the narration text. Specifically, for the audio duration corresponding to the narration text, if the predicted audio duration for a certain spatial segment is longer than the visual duration, non-factual modifiers in the narration text are deleted or short sentences are merged; if the predicted audio duration for a certain spatial segment is shorter than the visual duration, some detail observation statements strongly related to the visuals are inserted into the narration text, such as "Please pay attention,..."; if the total duration exceeds the initial video duration, the narration text can be compressed from non-core functional areas according to the word count ratio. For natural coherence, the first language model can automatically mark discontinuities (including logical jumps, stylistic abrupt changes, and rhythmic breaks), inserting transitional sentences before each discontinuity, such as "Next, we...", and converting passive voice and long relative clauses into short subject-verb-object structures, while appropriately inserting pause markers [PAUSE] before key information. For visual synchronization, the first language model performs detection to ensure that the narration text only uses verifiable visual elements from the visuals. For preliminary compliance screening, it can scan for absolute terms (such as "best," "top-level," etc.) and replace them with relative expressions (such as "better," "higher configuration"); it can identify promises (such as "guarantee," "assure," "must be," etc.) and convert them into conditional statements. As a result, the first language model outputs a second version of the narration text with multidimensional optimization to improve speech fluency and duration matching accuracy.
[0060] The third round is used for factual accuracy and compliance calibration reasoning. In the third round, the second version of the explanatory text is input into a pre-built third prompt word template to generate third prompt words. Guided by the third prompt words, the first language model corrects the numerical values involved in the second version of the explanatory text based on authoritative data sources, and completely removes commitment statements, outputting the final version of the explanatory text.
[0061] Optionally, within the thought chain reasoning process, Speech Synthesis Markup Language (SSML) can be generated, outputting a final corpus containing SSML control markers to ensure the linguistic expressiveness of the narration text aligns with the dynamic rhythm of the video. The SSML control markers can include time anchors for alignment with the video timeline, speech rate adjustment instructions, key information emphasis instructions, and silence / pause instructions. The parameter values of the SSML control markers can be dynamically determined based on the start and end times of the corresponding video frames, the motion state of the image, and the timbre characteristics.
[0062] For example, in the first round of inference, the first language model generates a preliminary explanatory corpus and its corresponding SSML structural framework based on the structured information obtained from video spatial segmentation and timbre parameters: an emphasis tag <emphasis> is added as an emphasis candidate at the core highlights to emphasize the core highlights of the property presented in the video; a silent pause tag <break time="?" / > is inserted at the spatial segment transitions to match the rhythm of the video movement; the starting sentence of each segment is placed in the speech rate adjustment tag <prosody rate="?" / > to dynamically adjust the speech rate according to the timbre characteristics and the video motion index; and a time anchor point <mark name="..." / > is inserted at the beginning of each spatial segment to achieve millisecond-level synchronization between the speech and the video frame.
[0063] In the second round of multidimensional quality collaborative optimization reasoning, the parameter values of SSML are dynamically fine-tuned by combining the pre-synthesized duration report and the video motion index: the accurate speech rate value is calculated and filled in based on the actual available duration; the duration of the silent segment is adjusted according to the duration deviation at the scene switching point; and the emphasis intensity is rhythmically adapted to avoid speech timeout due to overemphasis.
[0064] In the third round of reasoning, based on the precise start and end timestamps of the video frames corresponding to each spatial segment in the video, standard format time anchor marks are generated. At the same time, the syntactic validity of all SSML control marks is verified, invalid marks left over due to content deletion are removed, and explanatory text conforming to the SSML specification is output.
[0065] After generating the explanatory text, a Text-to-Speech (TTS) engine can be invoked. The explanatory text is input into the TTS engine to obtain the corresponding audio, which is then played synchronously with the video footage. By intelligently generating both the explanatory text and audio, the problem of inconsistent explanation quality due to varying levels of expertise among property providers can be avoided, effectively improving the overall quality of the explanations for the target property's video.
[0066] This application's embodiments achieve a high degree of alignment between the explanatory text and video visuals across four dimensions: semantics, temporality, rhythm, and emotion. The explanatory text is generated based on actually visible objects, key attributes, and property selling points in the video visuals, avoiding false descriptions and enhancing content credibility and user trust. Precise time anchors are embedded in the explanatory text, and a temporal alignment algorithm ensures that the speech rhythm matches the visual transitions, providing users with an immersive experience where what you hear is what you see. The speech rate and pauses are dynamically adjusted based on the video's movement; the speech is slowed down to highlight details during static presentations and sped up to maintain fluency during fast-paced cuts, ensuring a high degree of match between the speech performance and the video's emotional tone, thus improving viewing comfort.
[0067] In some implementations, before generating the explanatory text, a cover frame is selected from the initial video. The process of generating the explanatory text also includes inputting the semantic features of the cover frame as constraints into the first prompt word template so that the first sentence of the generated explanatory text is semantically consistent with the content of the cover image.
[0068] During the generation of the target property video, a frame that meets the requirements is selected from multiple target frames as the cover frame. A video cover is then generated based on this cover frame and stitched into the video header to achieve a better display effect. In this embodiment, the cover frame can be selected before generating the explanatory text, guiding the generation of the explanatory text.
[0069] For example, the selected cover frame is input into a multimodal large model, which outputs the descriptive text of the cover frame. Building upon the first round of inference, a semantic guidance mechanism for the cover is added: the descriptive text of the cover frame and the spatial segmentation results are input into a first prompt word template to generate a first prompt word. Guided by this prompt word, when generating the first sentence of the first version of the explanatory corpus, the large speech model parses the descriptive text of the cover frame, extracts the main visual elements, and uses a fixed sentence structure as the opening sentence (e.g., "What you see on the cover… is the core highlight of this property") for transitional guidance. Subsequent explanatory text is still generated based on the spatial segmentation results. When the explanation reaches the space corresponding to the cover frame in chronological order, it can also echo the cover content in detail. In this way, the first sentence of the explanatory text is generated semantically based on the cover image, establishing a connection between the cover and the video content. When users watch the video of the target property, the cover image and the first sentence of the explanation form a strong semantic echo, improving the first impression conversion rate without affecting the accurate explanation of the subsequent video content.
[0070] The cover frame can also be selected at any time after the explanatory text is generated, according to the embodiments of this application. The above is only one implementation method, and there is no limitation on the timing of the cover frame generation.
[0071] When selecting a cover frame, a large-scale visual understanding model can be used to understand the content of video frames and extract those that meet the requirements. For example, for each target frame, features can be extracted and scored from multiple dimensions, including spatial transparency, highlight matching degree, and conversion feature adaptation degree. A weighted fusion is then used to obtain a comprehensive score for each frame, and the frame with the highest score is selected as the cover frame for the target property video. Spatial transparency measures the visual and structural openness and connectivity of the space presented by the video frame. Highlight matching degree measures whether the content of the video frame is consistent with the aforementioned spatial highlights and property selling points. Conversion feature adaptation degree measures whether the video frame aligns with the psychological needs and key concerns of the target homebuyer.
[0072] For example, the visual understanding big model is based on multiple big model capability modules built-in or collaboratively invoked, including a semantic segmentation module (such as a finely tuned Mask2Former), a multimodal alignment module (such as CLIP-ViT-L / 14), and a lightweight visual recognition model, thereby achieving joint perception of spatial structure, semantic content, and user adaptation features.
[0073] In the spatial transparency dimension, the semantic segmentation module first performs high-precision pixel-level semantic segmentation on the target frame, accurately identifying key functional areas such as walls, doors, windows, living rooms, and balconies. Based on the segmentation results, it constructs a spatial topology map and analyzes the visibility and connectivity between functional areas. If two or more functional areas are simultaneously presented in a frame and there are no obstacles such as walls obstructing the line of sight, it is judged as highly transparent and assigned a high score. Otherwise, the score is exponentially decayed based on the number of visual fragmentations caused by obstacles, quantifying a continuous score reflecting the openness of the apartment layout. In the highlight matching dimension, for each highlight / selling point output by the spatial recognition model, the multimodal alignment module encodes the target frame into a visual feature vector and the text of the highlight / selling point into a semantic embedding vector. It then calculates the cosine similarity between the visual feature vector and the semantic embedding vector, obtaining multiple cosine similarities. The maximum cosine similarity is taken as the highlight matching score of the target frame. In the feature fit dimension, the lightweight visual recognition model activates a predefined visual feature rule base based on the target group label of the current property and drives a lightweight CNN classifier to perform fine-grained feature detection on the target frame. If a feature that highly matches the target group is detected, a positive incentive score is applied according to the rules. All three dimensions are normalized to scores in the interval [0,1].
[0074] By using a multi-dimensional integrated cover frame selection mechanism, we can select cover frames that are complete in space, have a wide field of view, better express the highlights of the property, and meet the preferences of the target audience. This enhances the informational power of the video cover, attracts user clicks, and improves conversion rates.
[0075] After selecting the cover frame, you can add key information that matches the core value of the property, or interactive visual elements such as animations and icons corresponding to that key information, to further enhance the cover's informational appeal. The placement, style, and color of these elements can be configured based on the visual information of the cover frame. For example, the cover frame and key information can be input into a data processing model. Guided by the textual semantics of the key information and the visual semantics of the cover frame, the model generates cover elements, displays these elements on top of the cover frame, and outputs the final cover.
[0076] Furthermore, this application embodiment also supports several optional auxiliary processes, including subtitle generation, background music generation, video ending generation, and explanation of the surrounding area of the community. Whether these processes are executed is dynamically determined based on the video content. In some implementations, if the first condition is met, subtitles are generated based on the explanatory text, and background music and an ending are configured. If the second condition is met, explanation of the surrounding area of the community is generated.
[0077] The first condition is used to assess whether the video as a whole is suitable for introducing enhanced elements. The purpose is to determine whether the current video, in terms of content structure, duration, or information carrying capacity, supports the addition of multimedia elements (such as subtitles, background music, end credits, etc.) without causing information interference or wasting resources. The first condition can be determined based on at least one of the video's content attributes, structural complexity, or playback duration.
[0078] For example, regarding subtitle generation, when the initial video length exceeds a preset threshold (e.g., 60 seconds), it indicates a large amount of information and potential for user distraction. In this case, the subtitle module is automatically activated, rendering the explanatory text into a clear and readable subtitle track to improve information reception efficiency and viewing experience. When the initial video length is less than the preset threshold (e.g., a quick preview video under 30 seconds), it is considered to have low information density, and subtitles can be omitted to avoid visual interference. Similarly, the introduction of background music is also related to the video length and rhythmic structure: longer videos typically contain multiple scene transitions or information modules, so it can be determined whether they have sufficient narrative rhythm to support background music. If the video structure supports this (e.g., including multiple segments such as exterior shots of a residential area, indoor panoramas, and close-ups of details), then light music with a coordinated style and controllable volume is automatically matched as an atmospheric complement; otherwise, no background music is added.
[0079] The second condition is used to determine whether the video is suitable for introducing explanations of the surrounding area of the community. Specifically, this can be done by detecting whether the initial video contains representations of the external environment, information gaps in the main narration text, and strategy signals driven by business tags. This comprehensive assessment determines whether the additional information (such as location-related explanations) can match the existing narrative logic, visual cues, or thematic focus of the video, thus avoiding semantic disconnect or content redundancy.
[0080] For example, if the explanatory text already covers the core selling points of the property (such as floor plan, decoration, and price) but does not mention its location value, an explanation of the surrounding area can be generated; if the explanatory text already mentions the location value, no explanation of the surrounding area will be generated. If the initial video contains visual elements such as exterior shots of the property, surrounding street scenes, or map animations, it is considered to have the contextual basis for expanding the explanation of the surrounding area, and an explanation of the surrounding area will be generated; if the initial video only focuses on indoor scenes and has no external environment scenes, no explanation of the surrounding area will be inserted. If the target property is tagged with a specified business tag, such as "school district property," an explanation of the surrounding area can be generated even if the video is short. Through the above conditional judgment mechanism that is deeply coupled with the content, structure, and duration of the video itself, intelligent and contextualized addition of optional processes is achieved, which not only avoids redundant processing but also effectively improves the information density and user appeal of the property video.
[0081] Specific details about the surrounding area of the residential complex can follow the property description. Based on the complex's location, information about its surrounding environment can be generated, including transportation, schools, hospitals, shopping malls, parks, etc. For example, leveraging high-precision geographic information data and a real-time rendering engine, a multi-dimensional living circle with a radius of 1 to 3 kilometers can be constructed centered on the complex. Interactive map rendering technology can dynamically present key supporting resources such as transportation networks, educational institutions, medical facilities, commercial complexes, and parks. Specifically, the complex's boundaries can be accurately marked on the map, followed by a heat map that visually displays the density distribution and service coverage of various facilities. For instance, a transportation facility heat map can show the accessibility of subway stations, bus hubs, and main roads; educational resource markers clearly indicate the distribution levels and school district information of key primary and secondary schools and kindergartens; and medical and commercial facilities are also distinguished using tiered icons and dynamic data labels. During the explanation, intelligent perspective switching is supported, automatically guiding viewers smoothly from a macro-level overview to micro-level facility details. It can dynamically simulate commuting times and travel mode preferences, transforming travel information into a visual timeline and route animation. It can also intelligently recommend preferred shopping, medical, and leisure routes, naturally integrating this information into the voice narration, forming a smooth and vivid narrative context. The entire explanation of the neighborhood relies on real-time rendering technology to achieve dynamic response and interactive feedback for map elements. Users can zoom, rotate, or click on specific facilities at any time to obtain details, thus gaining a comprehensive understanding of the quality of the living environment in an immersive experience, significantly enhancing the information support and scene immersion for home-buying decisions.
[0082] After using AI algorithms to perform a series of processes on the initial video, including spatial recognition and segmentation, explanatory text generation, explanatory audio generation, cover frame extraction, cover generation, subtitle generation, background music generation, video ending generation, and explanations of the surrounding area, the initial video can be used as a base to synthesize the spatial segmentation results, explanatory audio, cover, subtitles, background music, video ending, and explanations of the surrounding area to generate the target property video. The target property video features a spatial navigation, allowing users to see which spaces are included in the video while watching.
[0083] After generating the target property video, it can be published to the target application, allowing users to view the target property video using the target application on a second client.
[0084] Optionally, after the target property video is published, user interaction data can be recorded for subsequent video optimization and performance analysis. For example, the video content can be dynamically updated based on user questions and answers. The target application's video playback interface integrates an interactive Q&A module, allowing users to ask questions about property details in real time during viewing. Users can ask questions about content they want to know, and the user's questions, along with contextual information (such as the current playback time), are submitted to a second language model for semantic understanding and accurate answers, generating natural language responses that are presented to the user in real time.
[0085] In some optional embodiments, a backend video update module can be set up to update videos based on questions and answers, achieving dynamic updates of video content. The backend video update module can continuously collect and aggregate questions from different users and their corresponding model responses. Based on this aggregated data, it can trigger synchronous updates of video content: updating the explanatory text based on newly added or corrected key information, calling a TTS engine to generate the corresponding audio explanatory text, and synthesizing the updated audio explanatory text with the original video footage to generate an updated target property video. The updated target property video is pushed to the app for publication, replacing the original target property video and achieving a globally unified update. This ensures that when the next user watches the property video, the loaded video is the updated target property video, guaranteeing consistency across all users. The entire update process can be executed fully automatically on the server side without manual intervention, thereby achieving intelligent, dynamic, and globally synchronized updates of the explanatory content, significantly improving the timeliness and completeness of property information delivery.
[0086] Optionally, a full update can be performed on the entire video; alternatively, an incremental update can be performed based on spatial semantic anchors. This involves extracting key information from question-and-answer pairs from different users and accurately mapping it to existing structured spatial units within the video, while only correcting or enhancing relevant segments. During the video generation phase, spatial semantic anchors have been established for each spatial unit, including: a spatial category identifier; the original explanatory text segment corresponding to that space and its corresponding start and end timestamps in the video; and key items identified within that space and their attributes. These spatial semantic anchors serve as video metadata and are stored in the server-side database along with the video resources. The incremental update logic is as follows: the backend video update module calls the second major language model to perform spatial intent parsing on the question-and-answer pairs, determining whether the pairs point to a spatial unit of the target property (e.g., the master bedroom). If the question-and-answer pair can be mapped to a spatial unit, and the original explanatory text of that spatial unit does not explicitly or implicitly express the information of the question-and-answer pair, a structured update request is generated, containing the spatial category identifier, attribute fields to be supplemented / corrected, and recommended answer text. Subsequently, the backend video update module triggers the incremental script correction engine. This engine locates the original narration text fragment associated with the corresponding space and integrates the newly added semantic information into the original narration text. For example, the original narration text, "The master bedroom is equipped with a custom wardrobe and floor-to-ceiling windows," is updated to "The master bedroom is equipped with a custom wardrobe, floor-to-ceiling windows, and a south-facing bay window, with good natural light." The updated text fragment only replaces the original narration text corresponding to that space. Based on this, the corresponding audio fragment is resynthesized and seamlessly spliced back into the original video track using a timestamp alignment mechanism. Finally, the backend video update module generates a versioned incremental update package (containing the updated audio fragment, narration text fragment, and video version number) for subsequent distribution.
[0087] Users may ask questions related to the property's global attributes or external related information. The backend video update module will still call the second language model to combine the property's metadata to generate an accurate answer and display it to the user in real time. At this time, the Q&A information can be updated on the cover, or the community facilities explanation can be added, or a general information area can be reserved at the beginning / end of the video to insert new explanation segments.
[0088] The aforementioned backend video update module, as a backend microservice, can be deployed on the server side, such as a cloud server or backend cluster. The backend video update module communicates loosely with the target application client, video storage service, large language model inference service, and database service via API. After a user submits a question on the client, the question data is encrypted and transmitted to the backend video update module via the APP backend service. This module calls the large language model service to complete spatial intent parsing and answer generation, and then accesses the property metadata database to read the corresponding property metadata and video metadata. After completing local script corrections and audio synthesis, the backend video update module writes the newly generated audio clips and updated video metadata to object storage and updates the version index of the video resources. The next time the client loads the video of the target property, it automatically pulls the latest audio and video resources or incremental update packages through a version check mechanism, allowing the user to see the updated video of the target property.
[0089] In some alternative embodiments, when a user asks a question while watching the video explanation in the current space, the target application can update the explanation videos for other spaces that have not yet been viewed in real time based on the user's question and the answer information provided by the big model. The application will then provide explanations according to the new videos when the user navigates to other spaces, achieving personalized updates. Different users can see different explanation videos, providing a differentiated, context-aware, and personalized video experience. Personalized updates only apply to the current user's local cache, and the personalized video explanation content for different users is isolated from each other.
[0090] When initially publishing videos of target properties, the video resource file includes not only video resources and video metadata, but also a lightweight client-side video update module. This module enables terminal devices to regenerate and replace content locally, allowing for personalized, local updates without relying on a server.
[0091] When a user enters a space (such as the living room) and begins watching its explanation video, the interface simultaneously activates a Q&A input field. After the user asks a question, the client encrypts the question along with the user ID and uploads it to the server, allowing the server to call a large model to answer. After the server returns the answer, the client not only displays the answer on the current screen but also caches the question-and-answer pair in the local session context. The client's video update module iterates through the metadata of all spaces in the target property, determining whether the information in the question-and-answer pair points to a spatial unit of the target property and is explicitly or implicitly expressed in the original explanation text of the corresponding space (such as the balcony). If the original explanation text already covers the key information in the question-and-answer pair, it is considered duplicate information and is not updated. If the original explanation text does not cover the key information in the question-and-answer pair, the client's video update module rewrites the original explanation text for that space locally, integrating the semantics of the question-and-answer pair into the original explanation text to generate a personalized rewritten explanation text. It then calls the device's lightweight TTS engine to generate a new audio segment, replacing the local audio resource corresponding to that space. When the user subsequently navigates to the balcony, the player automatically loads the updated personalized explanation content. Since the entire update process is completed locally, there is no need to re-download the entire video, which ensures real-time performance and saves network bandwidth.
[0092] In this embodiment, the client-side video update module is implemented as a low-resource-consumption, distributeable with video packages, and policy-configurable intelligent client component. Its core functions include: parsing and storing metadata for each space; connecting to the server-side question-and-answer results; performing content coverage verification based on space metadata; rewriting the narration text for the corresponding space in a context-adaptive manner when update conditions are met; and calling local TTS to replace audio. The client-side video update module only interacts with the server once to obtain the answer when the user asks a question; all other logic is completed on the terminal. The server can dynamically control the behavior of the client-side video update module by issuing lightweight policy files (such as update trigger thresholds and allowed rewrite sentence templates), balancing flexibility and security. By delegating update decision-making and execution capabilities to the client and using the original narration text as the verification benchmark, the method provided in this embodiment can achieve efficient, personalized, and on-demand driven video content evolution while ensuring content accuracy, solving the current problems of untimely and cumbersome updates.
[0093] Optionally, to prevent distortion, logical inconsistencies, or deviations from the original facts caused by frequent or redundant updates, a local update limit and an automatic rollback mechanism can be set. When the client-side video update module performs each space-specific content replacement, it records the update operation log and increments a counter. If the number of local updates reaches a preset threshold (e.g., 3 times), or the cumulative text modification ratio exceeds a set ratio (e.g., 50%), the client-side video update module will stop subsequent update operations. Optionally, personalized update content can be saved in a temporary cache of the current session, with its lifecycle tied to the property browsing session. If the user exits the property page, switches to another property, or actively clears the application cache, an automatic cleanup process can be triggered, deleting the generated personalized audio and video clips and restoring the original explanation resources for each space.
[0094] Optionally, after each personalized update, the client-side video update module not only retains the original explanatory video resources but also persistently stores the latest generated personalized explanatory version locally (including rewritten text, synthesized audio, and corresponding timestamp mappings). This allows users to choose to view either the original explanatory video or the previously updated personalized version during subsequent visits. When a user re-enters the property, the target application detects the existence of historical update records and displays a selection control on the playback interface, providing the option to use the previously updated content or the original explanatory content. This selection control also supports two granularities: a global one-time selection and a space-specific selection. For example, a user can decide to continue using the personalized explanatory version for the master bedroom while reverting to the original version for the kitchen. The selection result is recorded locally and the corresponding audio and video resources are dynamically loaded when playing the corresponding space. Furthermore, users can switch between viewing the original and updated versions of a particular space at any time, enhancing transparency and trust. This allows highly interested homebuyers to focus on different dimensions of information during multiple browsing sessions and flexibly access historical personalized content, thereby improving decision-making efficiency and experience depth.
[0095] Optionally, after the target property video is published, the cover image can be dynamically updated based on user questions and answers. For example, the cover image can be updated when the video is republished, so that the next user or the next time a user watches the video will see the new cover image. Alternatively, the cover image can be dynamically updated periodically based on user browsing information; for example, if more users are interested in the master bedroom, the master bedroom can be used as the cover image, and if more users have viewed the living room, the living room can be used as the cover image. These two methods achieve global, unified updates, ensuring that all users see consistent video explanations. Furthermore, while updating explanation videos for other spaces in real time, the video cover image can also be updated to achieve personalized updates, allowing different users to see different video cover images. Personalized video cover images are stored and updated locally, avoiding server load.
[0096] In summary, this application embodiment can utilize AI technology to perform in-depth secondary processing on the initial video, automatically generating one or more post-production content such as explanatory audio, subtitles, cover, background music, ending, and supporting explanations, resulting in a high-quality target property video containing diverse content. The target property video is then automatically published, achieving intelligent processing throughout the entire process from original video processing and multi-dimensional content generation to automatic video publishing, significantly improving the efficiency and quality of property video generation. Furthermore, after the target property video is published, this application embodiment also provides a video update module, which can dynamically update the explanatory content and cover of the target property video based on user questions and answers, achieving automatic updates of the property video. The method provided by this application embodiment allows the property photographer to focus on shooting the original property video without needing to provide explanations while shooting, reducing the difficulty of shooting the original property video and improving video quality. Furthermore, by generating clear and content-rich explanatory audio, users can quickly and comprehensively understand the property information; the addition of subtitles, cover, ending, background music, and supporting explanations to the target property video enhances the user's viewing experience.
[0097] In some of the processes described in the above embodiments and accompanying drawings, multiple operations are included that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the terms "first," "second," etc., used herein are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0098] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 3 As shown, the electronic device includes a memory 31 and a processor 32.
[0099] Memory 31 is used to store computer programs and can be configured to store various other data to support operation on the computing platform. Examples of this data include instructions for any application or method operating on the computing platform, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0100] Processor 32, coupled to memory 31, executes a computer program in memory 31 to perform the following: acquiring an initial video of a target property captured in a photograph, the initial video comprising multiple video frames, the target property comprising multiple spatial regions; utilizing a spatial recognition model, based on the temporal correlation between video frames and structured prior knowledge of the indoor scene, performing object detection and spatial category reasoning on the video frames to obtain preliminary spatial segmentation results, and performing secondary fine-grained correction on the boundaries of the spatial regions based on the preliminary spatial segmentation results to obtain spatial recognition results, the spatial recognition results including the start and end times of the video frames corresponding to each spatial region, spatial category labels, and items contained in the space; utilizing a large-scale visual understanding model... The system selects a suitable frame from multiple video frames as the cover frame. Utilizing the first major language model, it uses items within each space as source material and limits the length of the explanatory text for each space to the start and end times of the corresponding video frames. Through a multi-round structured reasoning mechanism, it optimizes the constraints based on the intermediate results of the preceding reasoning output round by round, generating explanatory text aligned with the content of the initial video. The explanatory text is then input into a speech synthesis engine to obtain corresponding audio, which is used to synchronize with the initial video. Finally, the initial video, cover frame, and audio are synthesized to generate the target property video, which is then published to the target application for users to view.
[0101] Furthermore, such as Figure 3 As shown, the electronic device also includes other components such as a communication component 33, a display 34, a power supply component 35, and an audio component 36. Figure 3 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 3 The components shown. Additionally... Figure 3 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the work node. In this embodiment, the work node can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device, or a server-side device such as a conventional server, cloud server, or server array. If the work node in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 3 The components within the dashed box; if the working node in this embodiment is implemented as a server-side device such as a conventional server, cloud server, or server array, it may be omitted. Figure 3 The component within the dashed box.
[0102] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0103] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0104] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0105] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0106] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0107] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, digital video disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium.
[0108] Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above method embodiments.
[0109] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0110] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for processing a house video, characterized in that, The method comprises: acquiring an initial video obtained by photographing a target house source, the initial video comprising a plurality of video frames, and the target house source comprising a plurality of space regions; using a space recognition model, performing object detection and space category reasoning on the video frames based on the time sequence correlation between the video frames and the structured prior knowledge of the indoor scene, obtaining a preliminary space segmentation result, and performing secondary fine-grained correction on the boundaries of the space regions based on the preliminary space segmentation result to obtain a space recognition result, the space recognition result comprising the start and end time of the video frames corresponding to the identified region of each space, the space category label, and the objects contained in the space; using a visual understanding large model, selecting a frame meeting the requirements from the plurality of video frames as a cover frame; using a first large language model, taking the objects contained in each space as materials, limiting the length of the explanation text of each space by the start and end time of the video frames corresponding to the identified region of each space, and generating an explanation text aligned with the picture content of the initial video through a multi-round iterative reasoning mechanism based on the intermediate result output by the previous round of reasoning to constrain and optimize each round; inputting the explanation text into a speech synthesis engine to obtain corresponding explanation audio, which is used for synchronous playback with the picture of the initial video; synthesizing the initial video, the cover frame and the explanation audio to generate a target house source video, and publishing the target house source video to a target application for users to watch the target house source video.
2. The method of claim 1, wherein, The method comprises: acquiring an initial video obtained by photographing a target house source, the initial video comprising a plurality of video frames, and the target house source comprising a plurality of space regions; using a space recognition model, performing object detection and space category reasoning on the video frames based on the time sequence correlation between the video frames and the structured prior knowledge of the indoor scene, obtaining a preliminary space segmentation result, and performing secondary fine-grained correction on the boundaries of the space regions based on the preliminary space segmentation result to obtain a space recognition result, the space recognition result comprising: inputting the plurality of video frames into the space recognition model, and performing the following operations in the space recognition model: slicing the plurality of video frames according to a second-level time window, and screening a plurality of target frames meeting the requirements in each time window to form a target frame sequence; inputting the target frame sequence into an analysis engine of the space recognition model to enable the analysis engine to predict the space category to which each target frame belongs, and the prediction results of the plurality of target frames constituting a space prediction sequence; performing batch processing analysis on the space prediction sequence by using a sliding window mechanism to identify potential space switching points with changes in space category, using a plurality of heterogeneous space recognition models to respectively predict the space category of the video frames in the vicinity of the potential space switching points, and determining a candidate space switching point based on the voting result; in a preset time interval adjacent to the candidate space switching point, densely extracting a plurality of frames at a time granularity higher than a regular sampling frequency, and using the analysis engine to re-predict the space category of the plurality of frames, detecting the probability mutation amplitude between frames by first-order difference operation based on the prediction result, and determining the time point with the largest mutation amplitude as a space segmentation boundary.
3. The method of claim 2, wherein, inputting the target frame sequence into an analysis engine of the space recognition model, so that the analysis engine predicts a space category to which each target frame belongs, comprising: extracting features of each target frame through a shared visual encoder to obtain global semantic features and local region features of each frame; stacking the global semantic features and the local region features with space-time position encodings respectively, and performing cross-frame modeling through a temporal context fusion module to generate enhanced global features and enhanced local features; detecting categories and locations of items included in each target frame based on the enhanced local features; combining a preset item-space mapping knowledge base, dynamically weighting contributions of each item to space discrimination according to the detected categories and locations of the items through an attention mechanism, and outputting a space category prediction result of each target frame in combination with the enhanced global features.
4. The method of claim 2, wherein, The method further comprises: inputting multi-scale feature maps generated by the analysis engine in the feature extraction stage into a highlight attribute detection subnetwork of the space recognition model; performing multi-task feature reasoning related to space semantics by the highlight attribute detection subnetwork based on the multi-scale feature maps, wherein each task focuses on different space highlight attribute dimensions; outputting highlight attributes possessed by each target frame according to the reasoning result.
5. The method according to any one of claims 1 to 4, characterized in that, using a first large language model, taking items contained in each space as materials, limiting lengths of space explanation texts by start and end times of video frames corresponding to recognition regions of each space, and generating explanation texts aligned with picture contents of the initial video through a multi-round iterative reasoning mechanism, comprising: generating a first prompt word based on the space recognition result, taking items contained in each space as materials, limiting lengths of space explanation texts by start and end times of video frames corresponding to recognition regions of each space, and outputting a first version of explanation texts under the guidance of the first prompt word; generating a second prompt word based on the first version of explanation texts and a pre-synthesized duration report, taking items contained in each space as materials, limiting lengths of space explanation texts by start and end times of video frames corresponding to recognition regions of each space, and outputting a second version of explanation texts under the guidance of the second prompt word, wherein the modification includes at least one of the following: explanation duration, natural coherence, picture relevance, and compliance; generating a third prompt word based on the second version of explanation texts, taking items contained in each space as materials, and outputting explanation texts aligned with picture contents of the initial video under the guidance of the third prompt word, wherein the first large language model corrects numerical values involved in the second version of explanation texts according to authoritative data sources.
6. The method of claim 5, wherein, The method further comprises: adding speech synthesis markers during generation of the first version of explanation texts, including at least one of the following: adding an emphasis label at a highlight explanation of a target property, adding a silent pause label matching a picture motion rhythm at a space region switching point, adding a speech speed adjustment label at a start of explanation texts corresponding to each space region, and inserting a time anchor point at a start time of each space region; adjusting parameter values of the speech synthesis markers during generation of the second version of explanation texts.
7. The method according to claim 5 or 6, characterized in that, The method further comprises: inputting the semantic features of the cover frame as a constraint condition to the large voice model, so that the first sentence of the explanation text is consistent with the semantic content of the cover picture.
8. The method of claim 1, wherein, The method further comprises: generating auxiliary content when the first condition is met, the auxiliary content including at least one of the following: subtitles corresponding to the explanation text, background music, and an ending, the first condition being used to determine whether the initial video supports additional multimedia elements in terms of content structure, duration, or information carrying dimension; generating a surrounding area explanation when the second condition is met, the second condition being used to determine whether the initial video is suitable for introducing a surrounding area explanation.
9. The method of claim 1, wherein, The method further comprises: receiving a question from a user about the target property during the playback of the target property video, and sending the question to a second large language model to enable the second large language model to generate answer information based on the question and target property metadata; dynamically updating the explanation audio of the target property video based on the question and the answer information, and publishing the updated target property video to the target application.
10. The method of claim 9, wherein, The dynamic updating of the explanation audio of the target property video based on the question and the answer information comprises: performing spatial intent analysis on the question and answer pair formed by the question and the answer information, and when the question and answer pair points to any spatial area of the target property, determining whether the explanation text corresponding to the any spatial area expresses the semantics of the question and answer pair; if the explanation text corresponding to the any spatial area does not express the semantics of the question and answer pair, integrating the semantics of the question and answer pair into the explanation text corresponding to the any spatial area to obtain updated explanation text; dynamically updating the explanation audio of the target property video based on the updated explanation text.
11. The method according to claim 9 or 10, characterized in that, The dynamic updating is performed by a video updating module deployed on a server, and the video updating module is configured to update the target property video on the server. Alternatively, the dynamic updating is performed by a video updating module deployed on a client, and the video updating module is configured to update the target property video locally on the client.
12. An electronic device, comprising: comprises: a memory and a processor, wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the method of any one of claims 1 to 11.
13. A non-transitory machine-readable storage medium, comprising: The non-transitory machine-readable storage medium stores executable code, and when the executable code is executed by the processor of the electronic device, the processor executes the method of any one of claims 1 to 11.