House Video Processing Method, Device, Storage Medium and Program Product
By performing frame-by-frame and video clip-level scene recognition for house video streams, and using sliding window technology to determine the start frame and end frame of scene semantics, the problem of inaccurate room video scene recognition in the prior art is solved, and the effect of synchronizing audio explanation and video picture is achieved.
Patent Information
- Application Number
- CN202411426925.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-10-12
AI Technical Summary
When configuring audio explanations for housing videos, the scene recognition accuracy is low, resulting in the problem of out-synchronization of audio explanations and video images.
By receiving the video stream of the target house, the scene recognition model is used to identify the video frames frame by frame, extract keyframes, and perform sliding window processing with the video clip as granularity to identify whether the video clip and the keyframe belong to the same scene semantics, thereby accurately determining the start frame and end frame of the scene semantics.
The accuracy of recognition of scene semantics corresponding to each spatial area in the house video is improved, as well as the accuracy of recognition of video frames corresponding to each scene semantics, ensuring that the audio explanation is synchronized with the video picture.
Smart Images

Figure CN119299773B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing, and in particular, to a method, device, storage medium, and program product for processing house videos. Background Art
[0002] During the online house viewing process, users can select the housing units they want to know about and then view the videos of these housing units online. These videos include videos of each room in the housing unit, which facilitates users to view and understand the situation of the house more intuitively and realistically. This can not only meet the needs of users for online house viewing but also save the time cost of users for offline house viewing. To facilitate users to understand the housing unit information more quickly and comprehensively, during the process of shooting the housing unit video, the housing unit provider can synchronously record audio explanation information for each space area such as the master bedroom and bathroom while shooting, so as to provide a voice introduction to the information of each space area. However, the implementation cost and difficulty of synchronously recording voice explanation information while shooting are extremely high, and the requirements for the housing unit provider's housing unit explanation ability for shooting the housing unit video are also relatively high.
[0003] Therefore, a solution for configuring audio explanations for housing unit videos has emerged, that is, after obtaining the pre-shot housing unit video, through technical identification of the scenes in the housing unit video to identify each space area in the housing unit, and then configuring audio explanations for each space area. This method can reduce the implementation cost and difficulty of configuring audio explanations for housing unit videos and is not limited by the broker's explanation ability. However, if the scene recognition results for each space area in the housing unit video are not accurate enough, there will be a problem of out-of-sync between the audio explanation and the picture in the video. Therefore, how to accurately identify the scenes in the housing unit video is a technical problem to be solved urgently. Summary of the Invention
[0004] Multiple aspects of this application provide a method, device, storage medium, and program product for processing house videos to improve the accuracy of scene recognition for house videos.
[0005] An embodiment of the present application provides a method for processing a house video, including: receiving a video stream obtained by shooting a target house, where the video stream includes multiple video frames, the target house includes multiple spatial regions, and each spatial region has a corresponding scene semantics; inputting the multiple video frames into a scene recognition model to perform scene semantics recognition at the granularity of video frames, and obtaining the scene semantics to which each video frame belongs and the probability that it belongs to the scene semantics; determining consecutive video frames belonging to the same scene semantics according to the scene semantics to which each video frame belongs; determining key frames from the consecutive video frames belonging to the same scene semantics according to the probability that each video frame in the consecutive video frames belonging to the same scene semantics belongs to the scene semantics; for each key frame, use a sliding window to slide in two directions away from the key frame in turn starting from the key frame, input the video segment falling into the sliding window each time into the scene recognition model to recognize the probability that each video segment belongs to the same scene semantics as the key frame at the granularity of video segments, and stop sliding when it is continuously recognized that a specified number of video segments do not belong to the same scene semantics as the key frame according to the probability that each video segment belongs to the same scene semantics as the key frame; determine the start frame and end frame corresponding to the scene semantics to which the key frame belongs according to the positions where the sliding window stops sliding in the two directions.
[0006] An embodiment of the present application further provides an electronic device, including a memory and a processor, where the memory is used to store a computer program, and the processor is coupled to the memory and is used to execute the computer program to implement the steps in each method provided by the embodiments of the present application.
[0007] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it causes the processor to be able to implement the steps in the above-mentioned method.
[0008] An embodiment of the present application further provides a computer program product, which includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, it causes the processor to be able to implement the steps in the above-mentioned method embodiments.
[0009] In an embodiment of the present application, the video stream of the target house is frame-by-frame recognized for scene semantics at the granularity of video frames, and key frames are extracted based on the recognized scene semantics of each video frame and the probability that it belongs to the scene semantics; further, at the granularity of video segments, starting from any key frame, sliding is performed in two directions away from the key frame, and it is recognized whether the video segments falling into the sliding window each time belong to the same scene semantics as the key frame; and the sliding stops when a specified number of consecutive video segments are recognized not to belong to the same scene semantics as the key frame. In this way, the start frame and end frame of the scene semantics to which the key frame belongs can be accurately determined, thereby improving the accuracy of recognizing the scene semantics corresponding to each spatial region in the house video and the accuracy of recognizing the video frames corresponding to each scene semantics. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0011] Figure 1 is a schematic flowchart of a method for processing a house video provided by an exemplary embodiment of the present application;
[0012] Figures 2a - 2c is a schematic diagram of the output of a scene recognition model in another exemplary embodiment of the present application;
[0013] Figure 3 is a schematic flowchart of another method for processing a house video provided by an exemplary embodiment of the present application;
[0014] Figure 4 is a schematic diagram of the structure of an electronic device provided by another exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0016] It should be noted that in the embodiments of this application where user information is involved, the user information involved in the embodiments of this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject. Additionally, various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0017] Aiming at the technical problem of the low accuracy of scene recognition for housing videos in the prior art, in the embodiments of this application, the video stream of the target house is frame-by-frame recognized for scene semantics at the granularity of video frames, and key frames are extracted based on the recognized scene semantics of each video frame and the probability that it belongs to this scene semantics; further, at the granularity of video segments, starting from any key frame, sliding is performed in two directions away from this key frame, and it is determined whether the video segments falling into the sliding window each time belong to the same scene semantics as this key frame; and when a specified number of consecutive video segments are recognized as not belonging to the same scene semantics as this key frame, the sliding stops. In this way, the start frame and end frame of the scene semantics to which the key frame belongs can be accurately determined, thereby improving the accuracy of scene semantics recognition for each spatial region in the housing video and the accuracy of video frame recognition corresponding to each scene semantics.
[0018] Furthermore, when continuously recognizing whether a specified number of video segments are the same as the scene semantics to which the key frame belongs, a proof-by-contradiction model is introduced to perform consistency verification on the recognition results of the scene recognition model to correct inaccurate recognition results. During this process, item information is extracted from a specified number of consecutive video segments that are recognized as different from the scene semantics to which the key frame belongs; according to the proof-by-contradiction rules existing between the item information and the scene semantics to which the key frame belongs, scene consistency verification is performed on the specified number of video segments. If the scene semantics to which the specified number of video segments belong are the same as the scene semantics to which this key frame belongs, the sliding window continues to slide; if the scene semantics to which the specified number of video segments belong are different from the scene semantics to which this key frame belongs, the sliding window stops sliding.
[0019] Further, the concept of target scene semantics is introduced. Target scene semantics refers to scene semantics that include multiple sub - scene semantics. For the case where there are multiple key frames belonging to the target scene semantics, each key frame can determine a video segment, and multiple key frames correspond to multiple video segments. These video segments belong to the same target scene semantics. In the embodiments of the present application, based on a multi - modal recognition model, lighting features, visual field features, and area features are recognized to divide the multiple video segments corresponding to the target scene semantics into sub - scene semantics, and a sub - scene semantics corresponding to each video segment is obtained.
[0020] The following will, with reference to the accompanying drawings, elaborate on the technical solutions provided by each embodiment of the present application.
[0021] Figure 1 It is a schematic flowchart of a method for processing a house video provided by an exemplary embodiment of the present application. As Figure 1 shown, the method includes:
[0022] S101: Receive a video stream obtained by shooting a target house. The video stream includes multiple video frames, and the target house includes multiple spatial regions, and each spatial region has a corresponding scene semantics;
[0023] S102: Input the multiple video frames into a scene recognition model to perform scene semantics recognition at the granularity of video frames, and obtain the scene semantics to which each video frame belongs and the probability that it belongs to this scene semantics;
[0024] S103: According to the scene semantics to which each video frame belongs, determine consecutive video frames belonging to the same scene semantics; according to the probability that each video frame in the consecutive video frames belonging to the same scene semantics belongs to this scene semantics, determine key frames from the consecutive video frames belonging to the same scene semantics;
[0025] S104: For each key frame, use a sliding window to slide in two directions away from the key frame respectively starting from the key frame. Input the video segments falling into the sliding window each time into the scene recognition model to recognize the probability that each video segment belongs to the same scene semantics as the key frame at the granularity of video segments, and stop sliding when it is continuously recognized that a specified number of video segments do not belong to the same scene semantics as the key frame according to the probability that each video segment belongs to the same scene semantics as the key frame;
[0026] S105: According to the positions where the sliding window stops sliding in the two directions, determine the start frame and end frame corresponding to the scene semantics of the key frame.
[0027] In the embodiments of the present application, the target house includes multiple spatial regions, and different types of spatial regions correspond to different scene semantics. The corresponding scene semantics can be defined according to the functions of the spatial regions. For example, according to the functions of the spatial regions, they can be divided into scene semantics such as bedrooms, bathrooms, kitchens, and living rooms. Among them, there can be multiple spatial regions in the target house that belong to the target scene semantics, and the target scene semantics can be, but are not limited to, bedrooms, bathrooms, kitchens, etc. For example, the target house can include multiple spatial regions that belong to the scene semantics of the bathroom.
[0028] In this embodiment, shooting the target house can obtain a corresponding video stream, and the video stream includes multiple video frames sorted in time sequence on the time axis; among them, the multiple video frames are obtained by shooting different spatial regions of the target house, and the multiple video frames can cover multiple spatial regions. For example, some video frames are obtained by shooting spatial region A, some video frames are obtained by shooting spatial region B, and so on. Among them, the video frames obtained by shooting the same spatial region may be continuous or discontinuous. For example, bedroom A is shot twice. The first time, 10 video frames are obtained, and then other spatial regions such as the living room and balcony are shot. After that, bedroom A is shot again and 30 video frames are obtained. It can be seen that the video frames obtained by shooting bedroom A are not continuous, and there will be video frames corresponding to other spatial regions between the video frames obtained in the first shooting and the video frames obtained in the second shooting.
[0029] In this embodiment, the content provider can provide the video stream obtained by shooting the target house to a computer device as the video stream for subsequent scene semantics recognition. Among them, there is no restriction on the content provider. For example, it can be a real estate agent; or it can be a landlord; or it can also be other relevant personnel, etc.; and there is no limitation on the type of computer device. The computer device can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT (Internet of Things) device, or it can be a server device such as a conventional server, a cloud server, or a server array.
[0030] Further, input multiple video frames included in the video stream into the scene recognition model to recognize the scene semantics at the granularity of video frames. Among them, the scene recognition model is pre-trained using sample video streams of some houses. The sample video stream includes multiple sample video frames, and the scene semantics to which these sample video frames belong can cover various scene semantics in the house, such as the living room, bedroom, bathroom, and kitchen, etc. Furthermore, the sample video frames in the sample video stream can be used to train the scene recognition model to obtain the scene recognition model. The scene recognition model can be used to recognize the scene semantics of the video frames included in the video stream. For example, it can recognize that the scene semantics to which the video frame belongs are the living room, bedroom, bathroom, or kitchen, etc. The scene recognition model can be, but is not limited to, a MAE (Masked AutoEncoders) model.
[0031] Among them, for any input video frame, the scene recognition model will predict the probabilities that the video frame may belong to each scene semantics. For example, it can predict the probabilities that the video frame may belong to scene semantics such as the living room, bedroom, bathroom, and kitchen, etc., and select the first target probability that meets the first probability condition from the probabilities that the video frame may belong to each scene semantics. The scene semantics corresponding to the first target probability that is satisfied is used as the scene semantics to which the video frame belongs. For example, the first probability condition can be to select the highest probability from the probabilities that the video frame may belong to each scene semantics. The highest probability can be regarded as an example of the first target probability, but is not limited to this; the scene semantics corresponding to the highest probability is used as the scene semantics to which the video frame belongs. Furthermore, the scene recognition model outputs the scene semantics to which the video frame belongs and the probability that it belongs to this scene semantics.
[0032] As Figure 2a shown, it is an example of the scene recognition model outputting the scene semantics to which the video frame belongs and the probability that it belongs to this scene semantics. See Figure 2a , the horizontal axis is the duration of the video stream, and the vertical axis is the probability of the scene semantics to which the video frame belongs. In Figure 2a , the value range of the probability is from 0 to 1. 0 means it is impossible to belong to this scene semantics, 1 means it belongs to this scene semantics, and between 0 and 1 means the likelihood of belonging to this scene semantics. From 0 to 1, the likelihood gradually increases. It should be understood that Figure 2a this value range of the probability in Figure 2a is only an example and does not limit this embodiment. Further, Figure 2a different colors in Figures 2b - 2c represent different scene semantics. The scene semantics in
[0033] Further, determine consecutive video frames belonging to the same scene semantics according to the scene semantics to which each video frame belongs; determine key frames from the consecutive video frames belonging to the same scene semantics according to the probability of each video frame belonging to the scene semantics in the consecutive video frames belonging to the same scene semantics; in this process, select, from the consecutive video frames belonging to the same scene semantics, video frames with a second target probability that satisfies the second probability condition of belonging to the scene semantics as the key frames of the scene semantics; wherein, the second probability condition may be to select the video frame with the highest probability of belonging to the scene semantics from the consecutive video frames of the same scene semantics, but is not limited thereto. It should be understood that the first probability condition and the second probability condition may be the same or different, and no limitation is made thereto.
[0034] Continuing from the above Figure 2a Example, Figure 2b Give an example of key frames, which are key frame 1, key frame 2, and key frame 3 determined for the scene semantics of the bedroom. Among them, for the scene semantics of the bedroom, there are three segments of consecutive video frames, and key frames are extracted for each segment of consecutive video frames respectively to obtain key frame 1, key frame 2, and key frame 3. It can be seen that there can be multiple segments of consecutive video frames belonging to the same scene semantics, and the key frames of the scene semantics may be located in different segments of consecutive video frames.
[0035] Among them, taking any key frame as a reference point to determine the start frame position and end frame position of the scene semantics to which the key frame belongs, and according to the start frame position and end frame position corresponding to the key frame, determine the start frame and end frame corresponding to the scene semantics to which the key frame belongs, and further determine a video segment belonging to the scene semantics according to the start frame and end frame. Among them, the same scene semantics may have one or more key frames, and each key frame can be used to determine a video segment of the scene semantics to which it belongs, and the video segment is the same as the scene semantics to which the key frame belongs.
[0036] In this embodiment, the same sliding window processing is performed for each key frame. Taking any key frame as an example, the description is as follows. Taking any key frame as a reference point, the sliding window slides in two directions away from this key frame in turn; among them, one of the two directions is the time direction towards the time stamp earlier than this key frame, and the other is the time direction towards the time stamp later than this key frame. The order of sliding in the two directions is not limited; the video segments falling into the sliding window each time are input into the scene recognition model to identify the probability that each video segment belongs to the same scene semantics as this key frame at the granularity of the video segment; for each video segment, the scene recognition model will predict the probability that this video segment belongs to the same scene semantics as the key frame, and then the probability that the video segment belongs to the same scene semantics as this key frame can be used to identify whether the video segment and the key frame belong to the same scene semantics. For example, if the scene semantics of the key frame is a bedroom, the scene recognition model can predict the probability that this video segment belongs to the bedroom, and then the scene semantics of this video segment can be judged according to the probability that this video segment belongs to the bedroom.
[0037] In an alternative embodiment, for each video segment, the scene recognition model will predict the probabilities that this video segment may belong to various scene semantics. For example, it predicts the probabilities that this video segment may belong to scene semantics such as a living room, a bedroom, a bathroom, and a kitchen, and selects the one with the highest probability from the probabilities that this video segment may belong to various scene semantics, and takes the scene semantics corresponding to the highest probability as the scene semantics of this video segment; then, it judges whether the scene semantics of this video segment is the same as that of this key frame; if not, it obtains the probability that this video segment belongs to the same scene semantics as this key frame, and judges whether this video segment and the key frame belong to the same scene semantics according to this probability.
[0038] Among them, at the granularity of the video segment, in the process of sequentially performing scene semantics recognition on the video segments falling into the sliding window, it is to find video segments adjacent to this key frame and having the same scene semantics as this key frame from the key frame. In this process, a probability threshold can be set. If the probability that the video segment belongs to the same scene semantics as the key frame is greater than or equal to this probability threshold, it can be considered that the video segment and the key frame belong to the same scene semantics; otherwise, if the probability that the video segment belongs to the same scene semantics as the key frame is less than this probability threshold, it can be considered that the video segment and the key frame do not belong to the same scene semantics.
[0039] Among them, in the process of performing a sliding window process starting from a key frame to find video segments with the same scene semantics as the scene to which the key frame belongs, there may be sudden changes in the scene semantics of video segments within a few consecutive sliding windows, that is, the scene semantics of a few consecutive video segments are different from the scene semantics of the key frame. Among them, the multiple video segments that fall into the sliding window in sequence during the continuous sliding of the sliding window are consecutive video segments. In other words, consecutive video segments refer to multiple video frequency bands that are immediately adjacent in chronological order. However, a few consecutive video segments with sudden changes in scene semantics are not the true boundaries of the key frame, because when continuing to slide away from the key frame from these video segments with sudden changes in scene semantics, it is still possible to find video segments with the same scene semantics as the key frame. If these video segments with sudden changes in scene semantics are used as the "boundaries" of the key frame, video segments with the same scene semantics outside this "boundary" may be lost.
[0040] Based on this, in order to exclude the interference caused by sudden changes in the recognition results within a few consecutive sliding windows in the short term, in this embodiment, a specified number is determined in advance. The specified number can be determined according to empirical values. When the video segments within less than the specified number of consecutive sliding windows are different from the scene to which the key frame belongs, it is considered that the scene semantics of the video segments within less than the specified number of consecutive sliding windows have undergone sudden changes in this case. These video segments with sudden changes in scene semantics are not the true boundaries of the key frame. Therefore, the sliding window process can continue to identify video segments with the same scene semantics as the key frame. When it is continuously recognized that the scene semantics of the specified number of video segments are different from the scene semantics of the key frame, it can be considered that the boundary of the key frame has been found, and then the sliding window can stop sliding to stop the recognition operation of the video segments.
[0041] In this embodiment, the sliding window starts from any key frame and slides successively in two directions away from the key frame; among them, one of the two directions is the time direction with a time stamp earlier than that of the key frame, and the other direction is the time direction with a time stamp later than that of the key frame; for each direction, with the key frame as the reference point, the sliding window starts from the key frame and slides successively in the direction away from the key frame until the boundary of the key frame on the time axis in this direction is found, and the boundary is the time position on the time axis that does not belong to the same scene semantics as the key frame; when the boundaries are found in both directions, the sliding window stops sliding in both directions. Furthermore, according to the time positions where the sliding window stops sliding in the two directions, the start frame and the end frame corresponding to the scene semantics of the key frame are determined, where the start frame can be determined by the start position, and the start frame position is the farthest time position where the sliding window slides in the time direction with a time stamp earlier than that of the key frame; correspondingly, the end frame can be determined by the end frame position, and the end frame position is the farthest time position where the sliding window slides in the time direction with a time stamp later than that of the key frame; according to the start frame and the end frame, the video segment corresponding to the scene semantics of the key frame is determined, and the scene semantics of this video segment is the same as that of the key frame.
[0042] Continuing with the above Figure 2a example, as Figure 2c shown, Figure 2c is Figure 2a the effect diagram after the above house video processing method. It can be seen that Figure 2a the jitters and burrs have been removed. These burrs and jitters are the video frames that interfere with the scene semantics recognition of the video segment. For example, the burrs can be video frames that do not belong to the scene semantics different from that of a certain video segment, and the jitters can be the video frames where the scene is mis-switched due to the shaking of the mobile phone caused by human factors during the video stream shooting. After being processed by the above house video processing method, Figure 2c these burrs and jitters can be removed, and the start time and end time of the video segment corresponding to each scene semantics can be clearly seen, where the start time and end time are determined according to the start frame position and the end frame position.
[0043] In the embodiments of the present application, the video stream of the target house is frame-by-frame recognized for scene semantics at the granularity of video frames, and key frames are extracted based on the recognized scene semantics of each video frame and the probability that it belongs to the scene semantics; further, at the granularity of video segments, starting from any key frame, sliding is performed in two directions away from the key frame, and it is recognized whether the video segments falling into the sliding window each time belong to the same scene semantics as the key frame; and sliding stops when a specified number of consecutive video segments are recognized not to belong to the same scene semantics as the key frame. Thus, the start frame and end frame of the scene semantics to which the key frame belongs can be accurately determined, thereby improving the accuracy of recognizing the scene semantics corresponding to each spatial region in the house video and the accuracy of recognizing the video frames corresponding to each scene semantics.
[0044] Figure 3 FIG. is a schematic flowchart of another method for processing a house video provided by an exemplary embodiment of the present application. As Figure 3 shown, this process involves the terminal device 31 of the content provider and the computer device 30 that executes the house video processing method; among them, the terminal device 31 of the content provider and the computer device 30 can be the same device or different devices, and this is not limited. Moreover, the deployment location of the scene recognition model 32 is not limited. For example, it can be deployed on the computer device 30; or, it can also be deployed on the terminal device 31 of the content provider. Figure 3 In the example, the terminal device 31 of the content provider and the computer device 30 are different devices, and the scene recognition model 32 is deployed on the computer device 30 for example, but it is not limited to this.
[0045] In this embodiment, the content provider uploads the video stream obtained by shooting the target house to the computer device 30; after receiving the video stream, the computer device 30 inputs the multiple video frames included in the video stream into the scene recognition model 32. As Figure 3 shown in S301 in, the scene recognition model 32 recognizes the scene semantics at the granularity of video frames, and obtains the scene semantics to which each video frame belongs and the probability that it belongs to the scene semantics.
[0046] Optionally, during the process of the scene recognition model recognizing the scene semantics of the video frames, for any video frame, feature extraction is performed to obtain the feature information of the video frame; scene semantic prediction is performed on the video frame to obtain the probabilities that the video frame may belong to various scene semantics; a first target probability that meets the first probability condition is selected from the probabilities that the video frame may belong to various scene semantics, and the scene semantics corresponding to the first target probability and the first target probability are respectively used as the scene semantics to which the video frame belongs and the probability that it belongs to the scene semantics.
[0047] In an alternative embodiment, before inputting multiple video frames into the scene recognition model, the scene semantic range can also be determined in advance. The determination method includes: receiving the floor plan of the target house; determining the scene semantic range of the target house based on the floor plan of the target house; inputting the scene semantic range into the scene recognition model as a constraint condition for scene semantic recognition. Accordingly, when predicting the scene semantics of any video frame to obtain the probabilities that the video frame may belong to each scene semantic, it includes: predicting the scene semantics of any video frame within the scene semantic range to obtain the probabilities that the video frame may belong to each scene semantic. Since a limited scene semantic range is set, it is ensured that the output result of the scene recognition model is limited within this scene semantic range, thereby improving the recognition efficiency and accuracy of the scene semantics and the video frames corresponding to the scene semantics.
[0048] Further, as Figure 3 shown in S302, continuous video frames belonging to the same scene semantic are determined according to the scene semantic to which each video frame belongs; key frames are respectively determined from the continuous video frames belonging to the same scene semantic according to the probabilities that each video frame in the continuous video frames belonging to the same scene semantic belongs to this scene semantic; the same scene semantic may include at least one key frame. For the detailed content of this part, reference can be made to the foregoing embodiments and will not be elaborated here.
[0049] In an alternative embodiment, before using the sliding window to slide sequentially in two directions away from any key frame starting from the key frame, it further includes: obtaining the parameter information of the video stream, where the parameter information includes the duration and / or frame rate of the video stream; further, determining the width of the sliding window according to the duration and / or frame rate of the video stream; where the duration and frame rate of the video stream are positively correlated with the width of the sliding window. The longer the duration of the video stream, it means that during the shooting of the target house, the scene switching is slower, indicating that the amount of information contained in the video frames per unit time is more limited, so the width of the sliding window can be set larger; the higher the frame rate, it means that the number of video frames contained per unit time is more, so the width of the sliding window also needs to be set larger to improve the recognition efficiency of the overall video stream. Table 1 is given below, and Table 1 is a comparison table of the duration of a video stream and the width of the sliding window.
[0050] Table 1:
[0051] Video duration / minute Sliding window / second [0,2] 0.05s (2,5] 0.15s (5,10] 0.25s (10,+) 0.50s
[0052] As shown in Table 1, when the duration of the video stream is ≥ 0 minutes and ≤ 2 minutes, the width of the sliding window can be 0.05 seconds; when the duration of the video stream is > 2 minutes and ≤ 5 minutes, the width of the sliding window can be 0.15 seconds. Table 1 also includes the widths of the sliding windows corresponding to other video durations, which will not be elaborated here. Among them, considering that the sliding window has a preset width, both the starting frame position and the ending frame position are jointly determined by the number of times the sliding window slides in the corresponding direction and the preset width of the sliding window. For example, assume that the timestamp of a certain key frame is T seconds, the preset width of the sliding window is △t, and △t = 2 seconds; the sliding window slides N times in the time direction earlier than the timestamp of this key frame and then stops sliding, N ≥ 1 and N is a natural number; then the starting frame position = T - N × △t; correspondingly, the calculation method of the ending frame position can refer to that of the starting frame position, which will not be elaborated here.
[0053] Furthermore, as Figure 3 shown in S303, for each key frame, use the sliding window to slide in two directions away from the key frame in turn with a preset width, and input the video segment falling into the sliding window each time into the scene recognition model to identify the probability that each video segment belongs to the same scene semantics as the key frame with the video segment as the granularity. Optionally, the sliding window has a preset width, and the preset width can be determined according to the parameter information of the video stream, and the determination method can refer to the above embodiments; further, in the process of the scene recognition model performing scene semantics recognition with the video segment as the granularity, input the video segment falling into the sliding window each time into the scene recognition model, extract the features of each video frame in the video segment respectively, and obtain the feature information of each video frame in the video segment. Among them, the number of video frames included in the video segment is positively correlated with the preset width of the sliding window; furthermore, perform feature aggregation on the feature information of each video frame in the video segment to obtain an aggregated feature; predict the scene semantics of the aggregated feature to obtain the probability that the video segment belongs to the same scene semantics as the key frame, so as to be used to judge whether the video segment and the key frame belong to the same scene semantics.
[0054] In this embodiment, starting from any key frame, performing scene semantics recognition with the video segment as the granularity can not only significantly improve the recognition efficiency compared with the scheme of performing scene semantics recognition frame by frame, but also reduce the impact of the scene semantics mutation of a single video frame on the scene semantics recognition of the entire video segment, thereby improving the accuracy of scene semantics and the recognition of video frames corresponding to the scene semantics.
[0055] Optionally, in the process of continuously identifying that a specified number of video segments do not belong to the same scene semantics as any key frame according to the probability that each video segment belongs to the same scene semantics as the key frame, on the one hand, for the scene semantics to which any key frame belongs, a set probability threshold can be corresponding, and different scene semantics can correspond to the same or different set probability thresholds; furthermore, according to the probability that each video segment belongs to the same scene semantics as the key frame, video segments with an identification probability less than the set probability threshold can be identified; on the other hand, the video segments continuously identified as having a probability less than the set probability threshold are counted. When the count value for counting the number of video segments has not reached the specified number, each time one is identified, the count value is incremented by 1 until the specified number is reached; if the probability that the identified video segment belongs to the same scene semantics as the key frame is greater than the set probability threshold and the count value is not zero, the count value can be cleared, thereby ensuring that the count value is the video segments within the continuous sliding window.
[0056] In this embodiment, in the process of continuously identifying that the scene semantics of a specified number of video segments are different from the scene semantics to which the key frame belongs, as Figure 3 shown in S304, the anti - proof model and the object detection model are used to perform consistency verification on the scene semantics of the specified number of video segments to avoid misjudgment of the scene recognition model and improve the recognition rate of the scene recognition model. Among them, the specified number of video segments are sent into the object detection model, and at least one item information appearing in the specified number of video segments is obtained through the object detection model. Furthermore, it can be determined whether the at least one item information contains the target item information adapted to the scene semantics to which the key frame belongs; in the case where the target item information is not included in the at least one item information, the sliding window stops sliding. On the contrary, in the case where the target item information is included in the at least one item information, the sliding window can continue to slide. The following combines Figure 3 the steps S3041 and S3042 in to introduce how to perform consistency verification.
[0057] In an alternative embodiment, as Figure 3 shown in the step S3041, the specified number of video segments are input into the object detection model for item information detection to obtain at least one item information appearing in the specified number of video segments; the at least one item information and the scene semantics to which the key frame belongs are input into the anti - proof model. Different scene semantics correspond to their respective anti - proof rules, such as Figure 3As shown in step S3042, the disproof model verifies whether at least one piece of item information includes target item information adapted to the scene semantics to which the key frame belongs according to the disproof rules existing between the item information and the scene semantics; that the target item information is adapted to the scene semantics to which the key frame belongs means that the target item information does not violate facts or common sense under the scene semantics and can be used as a marker for the scene semantics and is the item information unique to the scene semantics. For example, if the scene semantics to which the key frame belongs is a bathroom, the target item information can be a toilet; if the scene semantics to which the key frame belongs is a kitchen, the target item information can be a range hood.
[0058] Among them, the disproof rules may include target item information; or, they may also include other item information other than the target item information. In the case where the disproof rules include target item information, the usage mode of the disproof rules is positive proof, that is, to verify whether at least one piece of item information includes the target item information set in the disproof rules. In the case of inclusion, it means that at least one piece of item information includes target item information adapted to the key frame scene semantics. On the contrary, it means that at least one piece of item information does not include target item information adapted to the scene semantics to which the key frame belongs; in the case where the disproof rules do not include target item information, the usage mode of the disproof rules is reverse proof, that is, to verify whether at least one piece of item information falls within the range of other item information included in the disproof rules other than the target item information; if at least one piece of item information falls within the range of other item information included in the disproof rules, it means that at least one piece of item information does not include target item information adapted to the scene semantics to which the key frame belongs. On the contrary, it means that at least one piece of item information includes target item information adapted to the scene semantics to which the key frame belongs. Whether it is positive proof or reverse proof, in this embodiment, it is to verify whether at least one piece of item information includes target item information adapted to the scene semantics to which the key frame belongs.
[0059] In this embodiment, in the case where at least one piece of item information does not include the target item information, it is determined that the scene semantics to which the specified number of continuously recognized video segments belong are not the same as the scene semantics to which the key frame does not belong; in the case where at least one piece of item information includes the target item information, it is determined that the scene semantics to which the specified number of continuously recognized video segments belong are the same as the scene semantics to which the key frame belongs. Further, the verification result is fed back to the scene recognition model for the scene recognition model to determine whether to continue with the sliding window processing. The sliding window processing can be implemented with reference to the above S303 step; further, as Figure 3As shown in step S305, when the verification result is the same, it means that the scene semantics of the specified number of consecutive video segments are the same as those of the key frame, and the sliding window process can continue; when the verification result is different, it means that the specified number of consecutive video segments and the key frame do not belong to the same scene semantics, and the sliding window process can be stopped. In one example, for instance, if the scene semantics of the key frame is the bathroom, the target item information adapted to the scene semantics of the bathroom can be a toilet; further, it can be verified whether the toilet is included in at least one item information. If not, it can be determined that the specified number of consecutive video segments do not belong to the scene semantics of the bathroom; if included, it is determined that the specified number of consecutive video segments belong to the scene semantics of the bathroom. Another example, if the scene semantics of the key frame is the kitchen, the target item information adapted to the scene semantics of the kitchen can be a range hood; further, it can be verified whether the range hood is included in at least one item information. If not, it can be determined that the specified number of consecutive video segments do not belong to the kitchen; if included, it is determined that the specified number of consecutive video segments belong to the kitchen.
[0060] In the embodiments of the present application, the concept of target scene semantics is introduced. The target scene semantics refers to the scene semantics that includes sub-scene semantics. For example, if the target scene semantics is the bedroom, it can include sub-scene semantics such as the master bedroom and the secondary bedroom; another example, if the target scene semantics is the balcony, it can include sub-scene semantics such as the living room balcony, the master bedroom balcony, and the secondary bedroom balcony. Among them, there are multiple key frames belonging to the target scene semantics. Each key frame has a corresponding start frame and end frame. According to the start frame and end frame, the corresponding video segment can be determined. Multiple key frames can determine multiple video segments. These video segments belong to the same target scene semantics, and each video segment has a corresponding sub-scene semantics.
[0061] First, in order to distinguish multiple video segments under the target scene semantics, a method of using labels to distinguish multiple video segments is proposed. For example, for the target scene semantics of the bedroom, when the video segment of the bedroom is recognized for the first time, it is labeled as Bedroom A; further, when the video segment of the bedroom is recognized for the second time, it is labeled as Bedroom B, so as to achieve the division of Bedroom A and Bedroom B. However, Bedroom A and Bedroom B are only simple label distinctions. Although it can be known that they both belong to the target scene semantics of the bedroom, it is impossible to identify which one is the master bedroom and which one is the secondary bedroom between Bedroom A and Bedroom B.
[0062] Therefore, in this embodiment, in order to divide multiple video segments of the target scene semantics into sub-scene semantics, this embodiment also provides an alternative implementation. In this alternative implementation, a multi-modal recognition model is used to divide the target scene semantics into sub-scene semantics. Among them, the multi-modal recognition model is pre-trained using sample video streams of some houses. The sample video stream includes multiple sample video frames, and these sample video frames include sample video frames of different sub-scene semantics under each target scene semantics. The multi-modal recognition model is trained using the sample video frames in the sample video stream to obtain the multi-modal recognition model. The multi-modal recognition model can perform lighting feature recognition, field-of-view feature recognition, and area size recognition on video frames, etc., to further divide the target scene semantics.
[0063] In this embodiment, according to the scene semantics to which each key frame belongs, it is determined whether there are multiple key frames that all belong to the target scene semantics. The target scene semantics refers to the scene semantics that includes multiple sub-scene semantics. If the determination result is yes, multiple video segments corresponding to the target scene semantics are determined. Each video segment is a continuous video frame determined according to the start frame and end frame corresponding to the scene semantics to which a key frame belongs. The multiple video segments of the target scene semantics are input into the multi-modal recognition model to identify the sub-scene semantics of the video segments of the target scene semantics, and the sub-scene semantics corresponding to each video segment are obtained. In this process, the multi-modal recognition model is used to perform at least one of the following operations:
[0064] Operation 1: Perform lighting feature recognition on each video segment to obtain the lighting feature information included in each video segment. The lighting feature information includes, but is not limited to, light intensity, light source direction, and number of light sources, etc. Further optionally, the natural light source and the artificial light source can be distinguished by combining the time information and the light source direction to further judge the lighting situation of the bedroom. For example, the master bedroom generally has better lighting than the secondary bedroom. In the case where it is recognized that bedroom A has better lighting than bedroom B, it can be considered that bedroom A is the master bedroom and bedroom B is the secondary bedroom.
[0065] Operation 2: Perform field-of-view feature recognition on each video segment to obtain the field-of-view feature information included in each video segment. The field-of-view feature information includes, but is not limited to, openness and field-of-view direction. For example, the master bedroom generally has a better field of view than the secondary bedroom. In the case where it is recognized that bedroom A has a better field of view than bedroom B, it can be considered that bedroom A is the master bedroom and bedroom B is the secondary bedroom.
[0066] Operation 3: Perform area feature recognition on each video segment to obtain the area feature information included in each video segment. The area feature information may include the area size. For example, the master bedroom generally has a larger area than the secondary bedroom. In the case where it is recognized that bedroom A has a better area than bedroom B, it can be considered that bedroom A is the master bedroom and bedroom B is the secondary bedroom.
[0067] For each video segment, identify the sub-scene semantics to which the video segment belongs according to at least one of the daylighting feature information, field of view feature information, and area feature information included in the video segment; where there is at least one video segment belonging to the same sub-scene semantics.
[0068] Further, as Figure 3 shown in S306, according to the positions where the sliding window stops sliding in two directions, determine the start frame and end frame corresponding to the scene semantics of the key frame. The start frame can be determined according to the start frame position, and the start frame position is the farthest time position where the sliding window slides in the time direction earlier than the time stamp of the key frame; correspondingly, the end frame can be determined according to the end frame position, and the end frame position is the farthest time position where the sliding window slides in the time direction later than the time stamp of the key frame; according to the start frame and end frame, determine the video segment corresponding to the scene semantics of the key frame, and the scene semantics of the video segment are the same as those of the key frame to which it belongs. Optionally, output the scene semantics recognition result in the form of a structured description file. For example, the output structure is as follows (unit / second):
[0069] {Start time: 0.0; End time: 30.3; Scene semantics: "Living room"};
[0070] {Start time: 30.3; End time: 75.6, Scene semantics: "Bedroom"};
[0071] {Start time: 75.6; End time: 95.0; Scene semantics: "Bathroom"};
[0072] {Start time: 95.0: End time: 133.6, Scene semantics: "Kitchen"}.
[0073] As shown in the above structure, each video segment corresponding to the scene semantics has a corresponding start time and end time, corresponding to the start frame position and end frame position respectively. For example, for the living room, the start time is 0.0, the end time is 30.3, and the duration is 30.3 seconds. The video frames during this period belong to the living room; the same applies to the subsequent scene semantics such as the bedroom, bathroom, and kitchen, which will not be elaborated here.
[0074] In this embodiment, based on the recognition result of the output scene semantics of the above scene recognition model, the generation model can generate corresponding audio explanation information. Among them, the generation model can determine the corresponding video segment of each scene semantics according to the recognition result of the scene semantics, that is, the scene semantics to which the key frame belongs and its corresponding start frame and end frame. Furthermore, based on the scene semantics to which each video segment belongs and the start frame and end frame of each video segment, audio explanation information adapted to each video segment is configured to obtain a target house video with audio explanation information, where the video segment and the adapted audio explanation information have the same start frame and end frame.
[0075] Optionally, for each video segment, it is identified whether the scene semantics to which the video segment belongs is the target scene semantics, and the target scene semantics refers to the scene semantics including multiple sub-scene semantics; if the judgment result is yes, the sub-scene semantics to which the video segment belongs are obtained, and based on the sub-scene semantics to which the video segment belongs and the start frame and end frame of the video segment, audio explanation information adapted to the video segment is configured to obtain a target house video with audio explanation information.
[0076] Optionally, the generation model can also receive text information, which is text information describing the video segment corresponding to the scene semantics and can be provided by the content provider. Based on the text information, each scene semantics, and the video segment belonging to the scene semantics, corresponding audio explanation information is generated. The audio content in the audio explanation information is generated according to the text information, and the audio explanation information has the same start time and end time as at least one video segment corresponding to the scene semantics.
[0077] In an alternative embodiment, the construction model can generate a three-dimensional house model of the target house based on each scene semantics and its corresponding at least one video segment. The process includes: obtaining initial data from at least one video segment corresponding to the same scene semantics; the initial data is 2D house type data corresponding to the scene semantics; furthermore, available plane information is extracted from the 2D house type data, and the plane information is processed to obtain three-dimensional space data of the spatial region corresponding to the scene semantics, so as to generate a three-dimensional space model corresponding to each spatial region in the target house according to the three-dimensional space data; the multiple three-dimensional space models corresponding to multiple spatial regions in the target house are integrated to generate an initial three-dimensional house model of the target house; furthermore, according to the image information in at least one video segment, corresponding textures and materials are mapped to different spatial regions of the initial three-dimensional house model to obtain the target three-dimensional house model of the target house.
[0078] Optionally, the audio explanation information associated with the video segment describing the spatial region and the three-dimensional spatial model corresponding to the spatial region can be merged to obtain a three-dimensional spatial model with audio explanation information; further, subsequent operations such as integrating the three-dimensional spatial model with audio explanation information are performed to obtain the target three-dimensional house model of the target house. Among them, the corresponding spatial region in the target house has audio explanation information, so that during the process of the user browsing the corresponding region in the target house model, the audio explanation information can be played, so that the user can better understand the house details.
[0079] The detailed implementation manners and beneficial effects of the steps in the method of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.
[0080] It should be noted that the execution subject of each step of the method provided in the foregoing embodiments can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps 101 to 103 can be device A; for another example, the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.
[0081] In addition, in some of the processes described in the foregoing embodiments and the accompanying drawings, multiple operations appear in a specific order, but it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as 101 and 102 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0082] Figure 4 This is a schematic structural diagram of an electronic device provided in another exemplary embodiment of the present application. As Figure 4 shown, the device includes: a memory 44 and a processor 45.
[0083] The memory 44 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0084] The memory 44 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0085] The processor 45, coupled to the memory 44, is configured to execute a computer program in the memory 44 for: receiving a video stream obtained by photographing a target house, the video stream including a plurality of video frames, the target house including a plurality of spatial regions, and each spatial region having a corresponding scene semantics; inputting the plurality of video frames into a scene recognition model to perform scene semantics recognition at the video frame level, to obtain the scene semantics to which each video frame belongs and the probability that it belongs to the scene semantics; determining consecutive video frames belonging to the same scene semantics according to the scene semantics to which each video frame belongs; determining key frames from the consecutive video frames belonging to the same scene semantics according to the probability that each video frame in the consecutive video frames belonging to the same scene semantics belongs to the scene semantics; for each key frame, using a sliding window to slide in two directions away from the key frame in turn starting from the key frame, inputting the video segment falling into the sliding window each time into the scene recognition model to recognize the probability that each video segment belongs to the same scene semantics as the key frame at the video segment level, and stopping sliding when it is continuously recognized that a specified number of video segments do not belong to the same scene semantics as the key frame according to the probability that each video segment belongs to the same scene semantics as the key frame; determining the start frame and the end frame corresponding to the scene semantics to which the key frame belongs according to the positions where the sliding window stops sliding in the two directions.
[0086] In an alternative embodiment, when the processor 45 inputs the plurality of video frames into a scene recognition model to perform scene semantics recognition at the video frame level, to obtain the scene semantics to which each video frame belongs and the probability that it belongs to the scene semantics, it is specifically configured to: input the plurality of video frames into a scene recognition model, perform feature extraction on any one of the video frames to obtain the feature information of the any one of the video frames; perform scene semantics prediction on the any one of the video frames to obtain the probabilities that the any one of the video frames may belong to various scene semantics; select a first target probability that meets a first probability condition from the probabilities that the any one of the video frames may belong to various scene semantics, and use the scene semantics corresponding to the first target probability and the first target probability as the scene semantics to which the corresponding any one of the video frames belongs and the probability that it belongs to the scene semantics, respectively.
[0087] In an alternative embodiment, before inputting the multiple video frames into the scene recognition model, the processor 45 is further configured to: receive the floor plan of the target house; determine the scene semantic range of the target house based on the floor plan of the target house; input the scene semantic range into the scene recognition model as a constraint condition for scene semantic recognition; accordingly, predict the scene semantics of any one of the video frames to obtain the probabilities that any one of the video frames may belong to each scene semantics, including: predicting the scene semantics of any one of the video frames within the scene semantic range to obtain the probabilities that any one of the video frames may belong to each scene semantics.
[0088] In an alternative embodiment, before the processor 45 starts to slide in two directions away from the key frame in turn using a sliding window starting from the key frame, it is configured to: obtain the parameter information of the video stream, where the parameter information includes the duration and / or frame rate of the video stream; determine the width of the sliding window according to the duration and / or frame rate of the video stream; wherein, the duration and / or frame rate of the video stream is positively correlated with the width of the sliding window.
[0089] In an alternative embodiment, when the processor 45 inputs the video segment falling into the sliding window each time into the scene recognition model to identify the probability that each video segment belongs to the same scene semantics as the key frame in terms of video segment granularity, it is specifically configured to: input the video segment falling into the sliding window each time into the scene recognition model, extract the feature information of each video frame in the video segment to obtain the feature information of each video frame in the video segment; perform feature aggregation on the feature information of each video frame in the video segment to obtain an aggregated feature; predict the probability that the video segment belongs to the same scene semantics as the key frame based on the aggregated feature.
[0090] In an alternative embodiment, when the processor 45 determines the start frame and the end frame corresponding to the scene semantics to which the key frame belongs according to the positions where the sliding window stops sliding in two directions, it is specifically configured to: when sliding in the time direction earlier than the time stamp of the key frame, use the farthest position of the sliding window from the key frame when it stops sliding as the start frame corresponding to the scene semantics to which the key frame belongs; when sliding in the time direction later than the time stamp of the key frame, use the farthest position of the sliding window from the key frame when it stops sliding as the end frame corresponding to the scene semantics to which the key frame belongs.
[0091] In an alternative embodiment, when the processor 45 continuously identifies that a specified number of video segments do not belong to the same scene semantics as the key frame according to the probability that each video segment belongs to the same scene semantics as the key frame, it is specifically configured to: according to the probability that each video segment belongs to the same scene semantics as the key frame, if a specified number of video segments with a probability less than a set probability threshold are continuously identified; then input the specified number of video segments into an object detection model for object detection to obtain at least one item information that appears in the specified number of video segments; input the at least one item information and the scene semantics to which the key frame belongs into a disproof model, and verify whether the at least one item information includes target item information adapted to the scene semantics to which the key frame belongs according to the disproof rule existing between the item information and the scene semantics; in the case that the at least one item information does not include the target item information, determine that the specified number of video segments do not belong to the same scene semantics as the key frame.
[0092] In an alternative embodiment, the processor 45 is further configured to: in the case that the at least one item information includes the target item information, determine that the specified number of video segments belong to the same scene semantics as the key frame, and continue to perform a sliding operation on the sliding window.
[0093] In an alternative embodiment, when the processor 45 continuously identifies a specified number of video segments with a probability less than a set probability threshold according to the probability that each video segment belongs to the same scene semantics as the key frame, it is specifically configured to: for each video segment, determine whether the probability that the video segment belongs to the same scene semantics as the key frame is less than the set probability threshold; if the determination result is less than the set probability threshold and the count value used to count the number of video segments has not reached the specified number, increment the count value by 1 until the specified number is reached; if the determination result is greater than or equal to the set probability threshold and the count value is not zero, clear the count value.
[0094] In an alternative embodiment, the processor 45 is further configured to: determine whether there are multiple key frames that all belong to a target scene semantics according to the scene semantics to which each key frame belongs, where the target scene semantics refers to a scene semantics including multiple sub-scene semantics; if the determination result is yes, determine multiple video segments corresponding to the target scene semantics, where each video segment is a continuous video frame determined according to the start frame and end frame corresponding to the scene semantics to which a key frame belongs; input the multiple video segments into a multi-modal recognition model, and use the multi-modal recognition model to perform at least one of the following operations: perform daylighting feature recognition on each video segment to obtain daylighting feature information included in each video segment; perform field-of-view feature recognition on each video segment to obtain field-of-view feature information included in each video segment; perform area size recognition on each video segment to obtain area feature information included in each video segment; for each video segment, recognize the sub-scene semantics to which the video segment belongs according to at least one of the daylighting feature information, field-of-view feature information, and area feature information included in the video segment; where there is at least one video segment belonging to the same sub-scene semantics.
[0095] In an alternative embodiment, the processor 45 is further configured to: determine a video segment corresponding to the scene semantics to which each key frame belongs according to the start frame and end frame corresponding to the scene semantics to which each key frame belongs; based on the scene semantics to which each video segment belongs and the start frame and end frame of each video segment, configure audio explanation information adapted to each video segment to obtain a target house video with audio explanation information.
[0096] In an alternative embodiment, when the processor 45 configures audio explanation information adapted to each video segment based on the scene semantics to which each video segment belongs and the start frame and end frame of each video segment to obtain a target house video with audio explanation information, it is specifically configured to: for each video segment, recognize whether the scene semantics to which the video segment belongs is a target scene semantics, where the target scene semantics refers to a scene semantics including multiple sub-scene semantics; if the determination result is yes, obtain the sub-scene semantics to which the video segment belongs, and based on the sub-scene semantics to which the video segment belongs and the start frame and end frame of the video segment, configure audio explanation information adapted to the video segment to obtain a target house video with audio explanation information.
[0097] Further, as Figure 4 shown, the electronic device further includes: other components such as a communication component 46, a display 47, a power supply component 48, and an audio component 49. Figure 4 Only some components are schematically shown, and it does not mean that the electronic device only includes Figure 4 the components shown. Additionally, Figure 4The components within the dashed-line box are optional components, rather than mandatory components, and can be determined according to the product form of the working node. The electronic device in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can be a server device such as a conventional server, a cloud server, or a server array. If the electronic device in this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it can include Figure 4 the components within the dashed-line box; if the working node in this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include Figure 4 the components within the dashed-line box.
[0098] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to be able to implement the steps in the above-mentioned method.
[0099] An embodiment of the present application further provides a computer program product, which includes computer programs / instructions that, when executed by a processor, cause the processor to be able to implement the steps in the above-mentioned method embodiments.
[0100] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0101] The above communication component is configured to facilitate communication, either wired or wireless, between the device where the communication component is located and other devices. The device where the communication component is located can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or combinations thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.
[0102] The above display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0103] The above power supply component provides power to various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing and distributing power for the device where the power supply component is located.
[0104] The above audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), which is configured to receive external audio signals when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode and a voice recognition mode. The received audio signals can be further stored in the memory or transmitted via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0105] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, compact disc read-only memory (CD-ROM), optical memory, etc.) that contain computer-usable program code.
[0106] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0109] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and a memory.
[0110] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0111] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0112] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0113] The above are only examples of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of the claims of the present application.
Claims
1. A housing video processing method, characterized in that: include: Receiving a video stream obtained by shooting a target house, the video stream including a plurality of video frames, the target house including a plurality of spatial regions, each of which has corresponding scene semantics; Inputting the multiple video frames into a scene recognition model to perform scene semantics recognition at the video frame granularity, and obtaining the scene semantics to which each video frame belongs and the probability of it belonging to the scene semantics; According to the scene semantics to which each video frame belongs, continuous video frames belonging to the same scene semantics are determined; according to the probability that each video frame in the continuous video frames belonging to the same scene semantics belongs to the scene semantics, key frames are determined from the continuous video frames belonging to the same scene semantics; For each key frame, a sliding window is used to slide in two directions away from the key frame in sequence starting from the key frame, and the video segments that fall into the sliding window each time are input into the scene recognition model to identify the probability that each video segment and the key frame belong to the same scene semantics based on the video segment granularity, and the sliding is stopped when a specified number of video segments are continuously identified as not belonging to the same scene semantics as the key frame according to the probability that each video segment and the key frame belong to the same scene semantics; According to the position of the sliding window when it stops sliding in two directions, the starting frame and the ending frame corresponding to the scene semantics to which the key frame belongs are determined.
2. The method according to claim 1, characterized in that Inputting the multiple video frames into a scene recognition model to perform scene semantics recognition at the video frame granularity, and obtaining the scene semantics to which each video frame belongs and the probability of it belonging to the scene semantics, including: Inputting the multiple video frames into a scene recognition model, and performing feature extraction on any video frame to obtain feature information of any video frame; Predicting scene semantics for any of the video frames to obtain probabilities that the any of the video frames may belong to each scene semantics; A first target probability that meets a first probability condition is selected from the probabilities that any of the video frames may belong to each of the scene semantics, and the scene semantics corresponding to the first target probability and the first target probability are respectively used as the scene semantics to which the corresponding any of the video frames belongs and the probability that it belongs to the scene semantics.
3. The method according to claim 2, characterized in that Before inputting the plurality of video frames into the scene recognition model, the method further includes: Receiving a floor plan of the target house; Determining a scene semantic range of the target house based on the floor plan of the target house; Inputting the scene semantic range into the scene recognition model as a constraint condition for scene semantic recognition; Accordingly, the scene semantics are predicted for any video frame to obtain the probability that any video frame may belong to each scene semantics, including: Within the scope of the scene semantics, the scene semantics of any video frame are predicted to obtain the probability that the any video frame may belong to each scene semantics.
4. The method according to claim 1, characterized in that: Before using the sliding window to sequentially slide in two directions away from the key frame starting from the key frame, the method further includes: Acquire parameter information of the video stream, wherein the parameter information includes duration and / or frame rate of the video stream; Determining the width of the sliding window according to the duration and / or frame rate of the video stream; The duration and / or frame rate of the video stream is positively correlated with the width of the sliding window.
5. The method according to claim 1, characterized in that: Inputting the video segments that fall into the sliding window each time into the scene recognition model to identify the probability that each video segment and the key frame belong to the same scene semantics with the video segment as the granularity includes: Inputting the video clips that fall into the sliding window each time into the scene recognition model, performing feature extraction on each video frame in the video clip, and obtaining feature information of each video frame in the video clip; Performing feature aggregation on feature information of each video frame in the video clip to obtain aggregated features; The probability that the video segment and the key frame belong to the same scene semantics is predicted based on the aggregated features.
6. The method according to claim 1, characterized in that Determining the start frame and the end frame corresponding to the scene semantics to which the key frame belongs according to the position where the sliding window stops sliding in two directions, including: When sliding in the time direction earlier than the timestamp of the key frame, the position of the sliding window farthest from the key frame when the sliding window stops sliding is used as the starting frame corresponding to the semantics of the scene to which the key frame belongs; When sliding in the time direction later than the timestamp of the key frame, the position of the sliding window farthest from the key frame when the sliding window stops sliding is used as the end frame corresponding to the scene semantics to which the key frame belongs.
7. The method according to claim 1, characterized in that According to the probability that each video segment and the key frame belong to the same scene semantics, continuously identifying a specified number of video segments and the key frame that do not belong to the same scene semantics, including: According to the probability that each video clip and the key frame belong to the same scene semantics, if a specified number of video clips with a probability less than a set probability threshold are continuously identified, the specified number of video clips are input into the object detection model for object detection to obtain at least one item information appearing in the specified number of video clips; Inputting the at least one item information and the scene semantics to which the key frame belongs into a disproof model, and verifying whether the at least one item information includes target item information that is compatible with the scene semantics to which the key frame belongs according to a disproof rule existing between the item information and the scene semantics; In a case where the at least one item information does not include target item information, it is determined that the specified number of video clips and the key frame do not belong to the same scene semantics.
8. The method according to claim 7, characterized in that Also includes: In a case where the at least one item information includes target item information, it is determined that the specified number of video clips and the key frame belong to the same scene semantics, and the sliding operation on the sliding window continues.
9. The method according to claim 7, characterized in that: According to the probability that each video segment and the key frame belong to the same scene semantics, a specified number of video segments whose probability is less than a set probability threshold are continuously identified, including: For each video segment, determining whether the probability that the video segment and the key frame belong to the same scene semantics is less than a set probability threshold; If the judgment result is less than the set probability threshold, and the count value used to count the number of video clips has not reached the specified number, then the count value is increased by 1 until it reaches the specified number; If the judgment result is greater than or equal to the set probability threshold, and the count value is not zero, the count value is cleared.
10. The method according to any one of claims 1 to 9, characterized in that: Also includes: According to the scene semantics to which each key frame belongs, determining whether there are multiple key frames that all belong to the target scene semantics, wherein the target scene semantics refers to the scene semantics including multiple sub-scene semantics; If the judgment result is yes, determine multiple video segments corresponding to the target scene semantics, each video segment is a continuous video frame determined according to a start frame and an end frame corresponding to the scene semantics to which a key frame belongs; Input the multiple video segments into a multimodal recognition model, and use the multimodal recognition model to perform at least one of the following operations: Perform lighting feature recognition on each video segment to obtain lighting feature information contained in each video segment; Performing visual field feature recognition on each video segment to obtain visual field feature information contained in each video segment; Identify the area size of each video segment to obtain area feature information contained in each video segment; For each video segment, the sub-scene semantics to which the video segment belongs is identified according to at least one of lighting feature information, field of view feature information and area feature information contained in the video segment; wherein there is at least one video segment belonging to the same sub-scene semantics.
11. The method according to any one of claims 1 to 9, characterized in that: Also includes: Determine, according to the starting frame and the ending frame corresponding to the scene semantics to which each key frame belongs, the video segment corresponding to the scene semantics to which each key frame belongs; Based on the scene semantics to which each video segment belongs and the start frame and the end frame of each video segment, each video segment is configured with audio explanation information adapted thereto, so as to obtain a target house video with audio explanation information.
12. The method according to claim 11, characterized in that Based on the scene semantics of each video segment and the start frame and end frame of each video segment, each video segment is configured with audio explanation information adapted thereto, so as to obtain a target house video with audio explanation information, including: For each video segment, identifying whether the scene semantics to which the video segment belongs is a target scene semantics, where the target scene semantics refers to a scene semantics including a plurality of sub-scene semantics; If the judgment result is yes, the sub-scene semantics to which the video segment belongs is obtained, and based on the sub-scene semantics to which the video segment belongs and the start frame and the end frame of the video segment, audio explanation information adapted to the video segment is configured to obtain a target house video with audio explanation information.
13. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor is coupled to the memory and is used to execute the computer program to implement the steps in the method according to any one of claims 1 to 12.
14. A computer-readable storage medium storing a computer program / instruction, characterized in that: When the computer program / instructions are executed by a processor, the processor is enabled to implement the steps of the method according to any one of claims 1 to 12.
15. A computer program product, characterized in that include: A computer program / instruction, when executed by a processor, causes the processor to implement the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Video classification method, device and equipment and computer readable storage medium
CN111274995A
Short video generation method and device, related equipment and medium
CN113453040A