System integrating video semantic segmentation and three-dimensional reconstruction technology

Through the system integrating video semantic segmentation and three-dimensional reconstruction technology, the problem of dynamic background changes in video and object recognition in complex environments is solved, real-time three-dimensional model generation and display, improving user experience and system efficiency, and promoting commercial application.

CN120264036APending Publication Date: 2025-07-04GUANGZHOU DONGLI SPORTS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510398907.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively deal with dynamic background changes in videos and object recognition in complex environments, resulting in low accuracy and efficiency of three-dimensional reconstruction, and lack of a convenient display platform, which affects user experience and commercial promotion.

Method used

A system that integrates video semantic segmentation and three-dimensional reconstruction technology, combined with the front-end uniapp framework and the back-end FastAPI framework, realizes real-time segmentation of video and synchronous reconstruction of three-dimensional models through the video semantic segmentation model and three-dimensional reconstruction model, and provides convenient upload and display processes.

Benefits of technology

It improves the efficiency of video processing and three-dimensional modeling, enhances the cross-platformity and user experience of the system, simplifies the operation process, and meets the needs of real-time and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264036A_ABST
    Figure CN120264036A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a video semantic segmentation and three-dimensional reconstruction technology integrated system and a front-end system, which are used for receiving a video uploaded by a user, marking and uploading the video and displaying a three-dimensional reconstruction module. The back-end system is used for receiving the video and the mark data uploaded by the front end and processing the video by applying a video semantic segmentation model; according to the method, the video uploaded by the user can be automatically processed, and the model with depth information and a three-dimensional structure is generated, so that the efficiency of video processing and three-dimensional modeling is greatly improved; the back-end system can send the video processing progress in real time and communicates with the front-end system through WebSocket, a user can know the video processing progress at any time in a real-time interaction mode, and the user experience is improved; through the arrangement of the uploading module, the marking module and the uploading marking module, a user can conveniently upload and mark a video, and the processes of video processing and three-dimensional modeling are greatly simplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and specifically to a system integrating video semantic segmentation and 3D reconstruction technology. Background Art

[0002] With the rapid development of 3D digital technology, the demand for generating high-quality 3D models based on video content has been growing in multiple fields such as e-commerce, health, and entertainment. Especially in e-commerce platforms, generating and displaying 3D models to enhance user experience and transaction conversion rate has become an important technology. However, there are still some obvious defects in the existing technical solutions in practical applications.

[0003] First of all, the existing semantic segmentation technologies mainly rely on static images or simple video frame processing, and it is difficult to effectively handle the dynamic background changes in videos and object recognition problems in complex environments, thus affecting the accuracy and efficiency of 3D reconstruction.

[0004] Secondly, most of the current technologies have not been able to achieve real-time segmentation of videos and synchronous reconstruction of 3D models. Especially when dealing with complex video content, it is difficult to meet the requirements of real-time and accuracy at the same time, which results in many application scenarios (such as virtual fitting, virtual reality, etc.) being unable to provide smooth 3D model generation and display, affecting the user experience.

[0005] Furthermore, although existing technologies can generate 3D models, there is a lack of a convenient and efficient display platform. Especially on lightweight platforms such as WeChat mini-programs, the existing solutions are often complex to operate, and it is difficult to enable users to easily upload videos and quickly generate and display 3D models. These defects limit the popularization and application of this technology in commercialization. Therefore, a system integrating video semantic segmentation and 3D reconstruction technology is proposed. Summary of the Invention

[0006] The object of the present invention is to provide a system integrating video semantic segmentation and 3D reconstruction technologies to solve the problems presented in the above background art. The existing semantic segmentation technologies mainly rely on static images or simple video frame processing, making it difficult to effectively handle dynamic background changes in videos and object recognition in complex environments, thus affecting the accuracy and efficiency of 3D reconstruction. Secondly, most of the current technologies have not been able to achieve real-time segmentation of videos and synchronous reconstruction of 3D models. Especially when dealing with complex video content, it is difficult to meet the requirements of real-time and accuracy simultaneously, which results in many application scenarios being unable to provide smooth 3D model generation and display, affecting the user experience. Moreover, although existing technologies can generate 3D models, there is a lack of a convenient and efficient display platform. Especially on lightweight platforms such as WeChat mini-programs, existing solutions are often complex to operate, making it difficult for users to easily upload videos and quickly generate and display 3D models. These defects limit the popularization and application of this technology in commercialization.

[0007] To achieve the above object, the present invention provides the following technical solution: A system integrating video semantic segmentation and 3D reconstruction technologies, including

[0008] A front-end system, which is used to receive videos uploaded by users, realize the marking and uploading of videos, and the display of the 3D reconstruction module;

[0009] A back-end system, which is used to receive the videos and marking data uploaded by the front-end, process the videos using a video semantic segmentation model, and return the processed video information to the front-end system.

[0010] Preferably, the above-mentioned front-end system uses the uniapp framework, and the back-end system uses the FastAPI framework.

[0011] Preferably, the above-mentioned front-end system includes an upload module, a marking module, and an upload marking module. The upload module is connected to the marking module, and the marking module is connected to the upload marking module;

[0012] The upload module is used to receive the videos uploaded by users and extract the first frame of the video as a marking image;

[0013] The marking module is used for users to mark foreground or background points on the first frame image of the video, or add marking frames;

[0014] The upload marking module is used for users to upload the marked videos and send the marking data and video information to the back-end system.

[0015] Preferably, the above-mentioned back-end system includes a video processing module and a 3D reconstruction module. The video processing module is connected to the 3D reconstruction module;

[0016] The video processing module is used to receive the marking data and video information, process the video using a video semantic segmentation model, and generate a mask;

[0017] The 3D reconstruction module is used to receive the processed video, run a 3D reconstruction model, and generate a 3D model.

[0018] Preferably, the above-mentioned video processing module is also responsible for sending the video processing progress in real time and communicating with the front-end system through WebSocket.

[0019] Preferably, the above-mentioned front-end system further includes a progress display module and a video display module, and the progress display module is connected to the video display module;

[0020] The progress display module is used to display the video processing progress in real time;

[0021] The video display module is used to display the processed video and the 3D reconstruction model.

[0022] Preferably, the above-mentioned video semantic segmentation model is the SAM2 model.

[0023] Preferably, the above-mentioned 3D reconstruction model is a built-in 3D reconstruction model.

[0024] Preferably, the above-mentioned upload marking module is connected to the video processing module, the video processing module is connected to the progress display module, and the 3D reconstruction module is connected to the video display module.

[0025] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects: By integrating video semantic segmentation and 3D reconstruction technologies, the present invention can automatically process the videos uploaded by users and generate models with depth information and 3D structures, thus greatly improving the efficiency of video processing and 3D modeling; The front-end system uses the uniapp framework and the back-end system uses the FastAPI framework, making the system have good cross-platform performance and scalability and being able to meet the needs and scenarios of different users; The back-end system can send the video processing progress in real time, communicate with the front-end system through WebSocket, and enable users to understand the video processing progress at any time in a real-time interaction manner, improving the user experience; Through the settings of the upload module, marking module, and upload marking module, users can conveniently upload videos and make marks, greatly simplifying the processes of video processing and 3D modeling. Description of the Drawings

[0026] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0027] Figure 1 It is a system schematic diagram of the present invention. Detailed implementation manners

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0029] It should be noted that the structures, ratios, sizes, etc. shown in the accompanying drawings of this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions that the present application can be implemented. Therefore, they do not have technical essence. Any modification of the structure, change of the proportional relationship or adjustment of the size, without affecting the effects that the present application can produce and the purposes that can be achieved, should still fall within the scope that the technical content disclosed in the present application can cover.

[0030] Embodiment

[0031] In the prior art, semantic segmentation technology mainly relies on static images or simple video frame processing, and it is difficult to effectively cope with the dynamic background changes in videos and object recognition problems in complex environments, thus affecting the accuracy and efficiency of three-dimensional reconstruction; secondly, most of the current technologies have not been able to achieve real-time segmentation of videos and synchronous reconstruction of three-dimensional models. Especially when dealing with complex video content, it is difficult to meet the requirements of real-time and accuracy at the same time. This has led to many application scenarios unable to provide smooth generation and display of three-dimensional models, affecting the user experience; furthermore, although existing technologies can generate three-dimensional models, there is a lack of a convenient and efficient display platform. Especially on lightweight platforms such as WeChat mini-programs, existing solutions are often complex to operate, and it is difficult to enable users to easily upload videos and quickly generate and display three-dimensional models. These defects limit the popularization and application of this technology in commercialization.

[0032] Please refer to Figure 1 , the present invention provides a technical solution: a system integrating video semantic segmentation and three-dimensional reconstruction technology, including

[0033] The front-end system is used to receive videos uploaded by users, implement video marking and uploading, and display the 3D reconstruction module;

[0034] The back-end system is used to receive the videos and marking data uploaded by the front-end, apply the video semantic segmentation model to process the videos, and return the processed video information to the front-end system.

[0035] The front-end system uses the uniapp framework, and the back-end system uses the FastAPI framework.

[0036] The front-end system includes an upload module, a marking module, and an upload marking module. The upload module is connected to the marking module, and the marking module is connected to the upload marking module;

[0037] The upload module is used to receive the videos uploaded by users and extract the first frame of the video as a marking image;

[0038] The steps for users to upload videos are as follows:

[0039] Users click on the "Upload Area" at the front-end;

[0040] A file selector pops up, and users select a video (from the camera or album);

[0041] The front-end calls the selectMedia() method and uses uni.chooseVideo to select the video;

[0042] After success, the video path is stored in mediaUrl, and the first frame of the video is extracted as imageUrl for marking;

[0043] Dynamically adjust the width and height of the container to maintain the aspect ratio of the video;

[0044] The marking module is used for users to mark foreground or background points on the first-frame image of the video, or add a marking box;

[0045] The steps for users to mark videos are as follows:

[0046] Users click the "Start Marking" button;

[0047] The front-end calls the startMarking() method, sets isMarking to true, and clears the previous marking points and boxes;

[0048] Users click on the first-frame image of the video to add foreground or background point markings, or click two points continuously to add a marking box;

[0049] The front-end records the relative coordinates (in percentage form) of the click positions and stores them in the marks and boxes arrays;

[0050] An upload marking module, which is used for users to upload the marked video and send the marking data and video information to the backend system;

[0051] The steps for users to upload the marked video are as follows:

[0052] The user clicks the "Start Processing" button, and the front end calls the uploadMask() method;

[0053] The front end checks whether there are valid marking points or boxes, and whether the video directory exists;

[0054] Use uni.request to send the marking data and video directory to the / predict endpoint of the backend;

[0055] Obtain the task ID and save the task information to local storage;

[0056] Establish a WebSocket connection and receive the processing progress in real time through the / ws / {task_id} endpoint.

[0057] The backend system includes a video processing module and a 3D reconstruction module. The video processing module is connected to the 3D reconstruction module, and the upload marking module is connected to the video processing module;

[0058] The video processing module is used to receive the marking data and video information, apply the video semantic segmentation model to process the video, and generate a mask. The video semantic segmentation model is the SAM2 model. The video processing module is also responsible for sending the video processing progress in real time and communicating with the front-end system through WebSocket;

[0059] The steps for the backend to process the video are as follows:

[0060] The / predict endpoint of FastAPI receives the marking data and video directory;

[0061] Generate a task ID and initialize the task status;

[0062] Use FFmpeg to split the video into frame images and save them to the output_videos1 directory;

[0063] Start a background task, use the SAM2 model to process each frame image, generate a mask, and apply the background color and reverse mode;

[0064] Synthesize the processed frame images into a video and save it to the output_videos3 directory;

[0065] During the processing, send progress updates in real time through WebSocket;

[0066] After processing is completed, a completion message is sent via WebSocket, including the URL of the processed video;

[0067] A 3D reconstruction module, which is used to receive the processed video, run a 3D reconstruction model, and generate a 3D model. The 3D reconstruction model is a built-in 3D reconstruction model;

[0068] The 3D reconstruction steps are as follows:

[0069] The user clicks on the "3D Reconstruction" module in the front end to enter the page. Clicking the upload button allows selection to upload from the processed video or the album and camera;

[0070] The uploaded video appears at the back end. At this time, the script is automatically loaded to run the built-in 3D reconstruction model.

[0071] The front-end system also includes a progress display module and a video display module, and the progress display module is connected to the video display module;

[0072] The progress display module is used to display the video processing progress in real time;

[0073] The steps for the front end to receive and display the processed video are as follows:

[0074] The front end receives progress updates via WebSocket and updates the progress bar and status information in real time;

[0075] When the completion message is received, the cacheAndPlayVideo(video_url) method is called to download and cache the processed video;

[0076] The path of the processed video is saved to the processedVideos array and displayed in the "Processed Video" area, where the user can click to play;

[0077] The video display module is used to display the processed video and the 3D reconstruction model.

[0078] The video processing module is connected to the progress display module, and the 3D reconstruction module is connected to the video display module. After successful 3D reconstruction, the model can be downloaded from the server to the front end and displayed.

[0079] Front-end and back-end data flow:

[0080] 1. Upload video:

[0081] User selects a video → The front end displays the video and the first frame image → The user marks the first frame image → The front end collects the marking data → The front end uploads the video file → The back end processes the video → The back end extracts video frames → Returns the URL of the first frame image and the video directory → The front end displays the first frame image for marking.

[0082] 2. Marked Data:

[0083] The user adds point or box marks at the front end → the front end records the relative coordinates and types → when uploading, the marked data is sent to the back end in the form of a JSON string.

[0084] 3. Video Processing:

[0085] The back end receives the marked data and the video directory → extracts video frames → generates masks using the SAM2 model → applies the reverse mode and background color → saves the processed frame images → synthesizes the processed video → returns the video URL.

[0086] 4. Real-time Progress Update:

[0087] The back end sends the processing progress in real time via WebSocket → the front end receives and updates the progress bar and status information.

[0088] 5. Display the Processed Video:

[0089] The front end receives the processed video URL → downloads and caches the video → displays it in the "Processed Video" area where the user can play it.

[0090] 6. 3D Reconstruction:

[0091] The front end clicks the "3D Reconstruction" button → uploads the video to the back end → the back end receives the video → runs the 3D reconstruction model → generates the model and the corresponding URL → the front end downloads and displays it.

[0092] 7. Display the 3D Model:

[0093] The front end downloads the model through the URL corresponding to the model → the front end uses the Three.js library to display the model.

[0094] In summary, by integrating video semantic segmentation and 3D reconstruction technologies, the present invention can automatically process the videos uploaded by users, generate models with depth information and 3D structures, thereby greatly improving the efficiency of video processing and 3D modeling; the front-end system uses the uniapp framework and the back-end system uses the FastAPI framework, making the system have good cross-platform performance and scalability, and can meet the needs and scenarios of different users; the back-end system can send the video processing progress in real time, communicates with the front-end system via WebSocket, and uses a real-time interaction method to enable users to understand the video processing progress at any time, improving the user experience; through the settings of the upload module, marking module and upload marking module, users can easily upload videos and make marks, greatly simplifying the processes of video processing and 3D modeling.

[0095] Those skilled in the art will understand that the features recited in the various embodiments and / or claims of the present invention can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and / and combined in various ways. All such combinations and / and combinations fall within the scope of the present invention.

Claims

1. A system integrating video semantic segmentation and 3D reconstruction technologies, characterized in that: including A front-end system for receiving videos uploaded by users, implementing video tagging and uploading, and displaying a 3D reconstruction module; A back-end system for receiving the videos and tagging data uploaded by the front-end, processing the videos using a video semantic segmentation model, and returning the processed video information to the front-end system.

2. The system integrating video semantic segmentation and 3D reconstruction technology according to claim 1, wherein: The front-end system uses the uniapp framework, and the back-end system uses the FastAPI framework.

3. The system integrating video semantic segmentation and 3D reconstruction technology according to claim 1, characterized in that: The front-end system includes an upload module, a tagging module, and an upload tagging module. The upload module is connected to the tagging module, and the tagging module is connected to the upload tagging module; The upload module is used to receive the videos uploaded by users and extract the first frame of the video as a tagging image; The tagging module is used for users to mark foreground or background points on the first frame image of the video, or add a tagging box; The upload tagging module is used for users to upload the tagged videos and send the tagging data and video information to the back-end system.

4. The system integrating video semantic segmentation and 3D reconstruction technology according to claim 3, characterized in that: The back-end system includes a video processing module and a 3D reconstruction module. The video processing module is connected to the 3D reconstruction module; The video processing module is used to receive the tagging data and video information, process the videos using a video semantic segmentation model, and generate a mask; The 3D reconstruction module is used to receive the processed videos, run a 3D reconstruction model, and generate a 3D model.

5. The system integrating video semantic segmentation and 3D reconstruction technology according to claim 4, characterized in that: The video processing module is also responsible for sending the video processing progress in real time and communicating with the front-end system through WebSocket.

6. The system integrating video semantic segmentation and 3D reconstruction technology according to claim 4, characterized in that: The front-end system further includes a progress display module and a video display module. The progress display module is connected to the video display module; The progress display module is used to display the video processing progress in real time; The video display module is used to display the processed videos and the 3D reconstruction model.

7. A system integrating video semantic segmentation and 3D reconstruction technologies according to claim 4, characterized in that: The video semantic segmentation model is the SAM2 model.

8. The system integrating video semantic segmentation and 3D reconstruction technology according to claim 4, characterized in that: The 3D reconstruction model is a built-in 3D reconstruction model.

9. The system integrating video semantic segmentation and 3D reconstruction technology according to claim 6, characterized in that: The upload tagging module is connected to the video processing module, the video processing module is connected to the progress display module, and the 3D reconstruction module is connected to the video display module.