Method, apparatus, device, and storage medium for generating remote training videos
Through image segmentation and synthesis of remote training videos, the problem of insufficient interactivity and fun in remote training is solved, real-time interaction between lecturers and students and accurate communication of training content is achieved.
Patent Information
- Application Number
- CN202010043553.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-01-15
AI Technical Summary
The existing remote training video display is too single, lacking interaction and fun, and the instructor's body language and expressions cannot be observed by students in real time, resulting in poor interaction.
By obtaining the video stream during the training process, using the image segmentation model to segment the instructor's independent portrait picture, teaching courseware picture and student picture, extract element content and location information, build a picture frame for the virtual lecture hall, and add element content to the corresponding location to generate AI lecture hall video.
It enables lecturers and students to observe the teaching status in real time, improves interactivity and fun, and quickly and accurately conveys training content.
Smart Images

Figure CN111242962B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a method, device, equipment and computer-readable storage medium for generating a remote training video. Background Art
[0002] Multimedia is the integration of multiple media, generally including various media forms such as text, sound and images. Multimedia is the embodiment of modern informatization and also the trend of social development. Especially in the field of education, multimedia education is also a part of modern informatization. Vigorously promoting multimedia education has become the trend of educational development, and at the same time, it makes up for the deficiencies in traditional teaching and can realize simultaneous teaching between different regions.
[0003] In order to solve the problem of geographical restrictions in traditional education methods, in the prior art, remote teaching using multimedia technology is realized through the Internet. However, in current remote teaching and other trainings, the lecturer needs to share courseware (PPT / or other documents) and the computer desktop with students. But when the students in class look at the computer, they cannot see the lecturer's expressions, movements and body languages. Therefore, the body language information of the teacher will be missed. In addition, its interactivity and interest are also relatively poor. Summary of the Invention
[0004] The main purpose of the present invention is to provide a method, device, equipment and computer-readable storage medium for generating a remote training video, aiming to solve the technical problem that the display of existing remote training videos is too single and lacks interactivity and interest.
[0005] To achieve the above object, the present invention provides a method for generating a remote training video, which is applied to a remote training platform. The method for generating the remote training video includes the following steps:
[0006] Obtain a video stream during the training process, where the video stream includes: a teaching video when the lecturer teaches and / or an interactive video when the trainee participates in the training;
[0007] Perform segmentation processing on the video stream through a preset image segmentation model to obtain independent pictures, where the independent pictures include one or more of an independent portrait picture of the lecturer teaching, an independent courseware picture for teaching, and an independent trainee picture;
[0008] Extract the element content in the independent pictures and the position information of the element content in the pictures;
[0009] Construct a picture frame of a virtual lecture hall according to the position information, and the picture frame is a picture layout for accommodating the independent pictures;
[0010] Add the element content to the corresponding position of the screen frame to obtain the training video of the AI classroom.
[0011] Optionally, after the step of obtaining the video stream during the training process, the following steps are further included:
[0012] Detect whether the teaching video in the video stream is a mixed video, where the mixed video includes the portrait video of the instructor and the teaching courseware video;
[0013] If the teaching video is a mixed video, use the portrait extraction algorithm to extract the independent portrait picture of the instructor in the face video, and use the character detection algorithm to extract the courseware information currently used by the instructor in the teaching courseware video to obtain the independent teaching courseware picture;
[0014] If the teaching video is a non-mixed video, execute the step of segmenting the video stream through a preset image segmentation model to obtain independent pictures.
[0015] Optionally, the step of segmenting the video stream through a preset image segmentation model to obtain independent pictures includes:
[0016] When the video stream is the teaching video of the instructor, calculate the depth value of the depth of field of each picture in the teaching video according to the preset depth of field formula;
[0017] Identify the foreground area and the background area of the picture in the teaching video according to the depth value, where the portrait picture is included in the foreground area;
[0018] Use the image matting algorithm to extract the foreground area from the teaching video to obtain the foreground video picture, and extract the background area from the teaching video to obtain the independent teaching courseware picture;
[0019] Identify the portrait picture in the foreground area according to the preset face recognition algorithm, and extract the portrait picture from the foreground area to form the independent portrait picture.
[0020] Optionally, the step of segmenting the video stream through a preset image segmentation model to obtain independent pictures includes:
[0021] When the video stream is the video of the students' listening and interaction, use face recognition technology to identify whether there are students in the listening and interaction video who meet the classroom interaction postures, where the classroom interaction postures include standing and raising hands;
[0022] If it exists, the human body of the student is scanned through a camera to obtain the portrait contour of the student, and the depth of field value of the portrait contour in the listening and interaction video is calculated according to a preset depth of field formula;
[0023] Using the depth of field value as the picture cutting critical point, all the pictures on the critical point in the listening and interaction video are cut out to form the independent student picture.
[0024] Optionally, the step of extracting the element content in the independent portrait picture, the independent teaching courseware picture and / or the independent student picture, and the position information of the element content in the picture includes:
[0025] Create a canvas according to the length and width of the independent picture, and select any corner point in the canvas as the coordinate origin to establish a two-dimensional coordinate system;
[0026] Based on the two-dimensional coordinate system, calculate the coordinate information of the lecturer's portrait or the student's portrait in the independent picture, and calculate the coordinate information of the teaching content of the teaching courseware in the independent picture;
[0027] Extract the portrait and courseware content from the independent picture according to the coordinate information.
[0028] Optionally, the step of constructing the picture frame of the virtual lecture hall according to the position information includes:
[0029] Use the independent teaching courseware picture as the background canvas of the AI lecture hall, and construct a coordinate system on the background canvas;
[0030] According to the coordinate information, draw a portrait filling area with the same shape as the portrait on the background canvas to obtain the picture frame;
[0031] The step of adding the element content to the corresponding position of the picture frame to obtain the training video of the AI lecture hall includes:
[0032] Fill the extracted portrait into the corresponding portrait filling area, and fuse the portrait with the background canvas through the boundary interpolation background fusion algorithm to obtain the training video.
[0033] Optionally, the depth of field calculation formula is:
[0034] ,
[0035] where δ is the allowable circle of confusion diameter, f is the lens focal length, F is the shooting aperture value of the lens, and L is the focusing distance.
[0036] To solve the above problems, the present invention also provides a device for generating remote training videos, which includes:
[0037] An acquisition module for obtaining a video stream during the training process, where the video stream includes: a teaching video when the instructor is teaching and / or an interactive video of the students during the training;
[0038] A segmentation module for segmenting the video stream through a preset image segmentation model to obtain independent images, where the independent images include one or more of an independent portrait image of the instructor, an independent teaching material image, and an independent student image;
[0039] An extraction module for extracting the element content in the independent images and the position information of the element content in the images;
[0040] A synthesis module for constructing a frame of a virtual lecture hall according to the position information, where the frame of the virtual lecture hall is a layout for accommodating the independent images; adding the element content to the corresponding positions of the frame of the virtual lecture hall to obtain a training video of an AI lecture hall.
[0041] Optionally, the device for generating remote training videos further includes a detection module for detecting whether the teaching video in the video stream is a mixed video, where the mixed video includes a portrait video of the instructor and a teaching material video;
[0042] If the teaching video is a mixed video, the segmentation module uses a portrait extraction algorithm to extract an independent portrait image of the instructor in the face video, and uses a character detection algorithm to extract the teaching material information currently used by the instructor in the teaching material video to obtain the independent teaching material image;
[0043] If the teaching video is a non-mixed video, the segmentation module is controlled to execute the step of segmenting the video stream through a preset image segmentation model to obtain independent images.
[0044] Optionally, the segmentation module includes a first calculation unit, an identification unit, and a first cutting unit, where:
[0045] The first calculation unit is used to calculate the depth value of the depth of field of each image in the teaching video according to a preset depth of field formula;
[0046] The identification unit is used to identify the foreground area and the background area of the image in the teaching video according to the depth value, where the portrait image is included in the foreground area;
[0047] The first cutting unit is used to extract the foreground area from the teaching video by using an image matting algorithm to obtain a foreground video frame, and extract the background area from the teaching video to obtain the independent teaching courseware frame; and according to a preset face recognition algorithm, identify the face frame in the foreground area, and extract the face frame from the foreground area to form the independent face frame.
[0048] Optionally, the segmentation module includes a face recognition unit, a second calculation unit, and a second cutting unit, where:
[0049] The face recognition unit is used to use face recognition technology to identify whether there are students in the listening and interaction video who meet the classroom interaction postures, where the classroom interaction postures include standing and raising hands;
[0050] If there are, the second calculation unit performs a human body scan on the student through a camera to obtain the human contour of the student, and calculates the depth of field value of the human contour in the listening and interaction video according to a preset depth of field formula;
[0051] The second cutting unit is used to use the depth of field value as the picture cutting critical point to cut out all the pictures in the listening and interaction video located on the critical point to form the independent student frame.
[0052] Optionally, the extraction module is used to create a canvas according to the length and width of the independent frame, and select any corner point in the canvas as the coordinate origin to establish a two-dimensional coordinate system; based on the two-dimensional coordinate system, calculate the coordinate information of the lecturer's portrait or the student's portrait located in the independent frame, and calculate the coordinate information of the teaching content of the teaching courseware located in the independent frame; according to the coordinate information, extract the portrait and courseware content from the independent frame.
[0053] Optionally, the synthesis module is used to use the independent teaching courseware frame as the background canvas of the AI lecture hall and construct a coordinate system on the background canvas; according to the coordinate information, draw a portrait filling area with the same shape as the portrait on the background canvas to obtain the picture frame; the step of adding the element content to the corresponding position of the picture frame to obtain the training video of the AI lecture hall includes: filling the extracted portrait into the corresponding portrait filling area, and fusing the portrait with the background canvas through a boundary interpolation background fusion algorithm to obtain the training video.
[0054] Optionally, the depth of field calculation formula is:
[0055] ,
[0056] Among them, δ is the diameter of the permissible circle of confusion, f is the focal length of the lens, F is the shooting aperture value of the lens, and L is the focusing distance.
[0057] In addition, to achieve the above-mentioned purpose, the present invention also provides a training device, which includes: a memory, a processor, and a remote training video generation program stored in the memory and run on the processor. When the remote training video generation program is executed by the processor, the steps of the remote training video generation method as described in any one of the above items are implemented.
[0058] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a remote training video generation program is stored. When the remote training video generation program is executed by the processor, the steps of the remote training video generation method as described in any of the above items are implemented.
[0059] The present invention provides a method for generating remote training videos, which mainly collects lecturers' teaching videos and students' training videos in real time, segments the videos based on an image segmentation model, extracts separate pictures, and extracts element content and its position information in the original pictures from the pictures, constructs a picture frame of the synthetic video according to the position information, and synthesizes the video stream into a new training video based on the picture frame to obtain an AI lecture video. Such a video allows both parties to observe the status of teaching and listening in real time, as well as improves interactivity, and can quickly and accurately convey the content of the training. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a schematic diagram of the structure of the operating environment of the remote training platform involved in the embodiment of the present invention;
[0061] Figure 2 A schematic diagram of a flow chart of an embodiment of a method for generating a remote training video provided by the present invention;
[0062] Figure 3 A schematic diagram of a flow chart of another embodiment of a method for generating a remote training video provided by the present invention;
[0063] Figure 4 A schematic diagram of the process of extracting element content provided by the present invention;
[0064] Figure 5 A schematic diagram of the flow chart of the screen layout of the classroom provided by the present invention;
[0065] Figure 6 A schematic flow chart of another embodiment of the method for generating a remote training video provided by the present invention;
[0066] Figure 7Schematic diagram of the functional modules of an embodiment of the apparatus for generating remote training videos provided by the present invention.
[0067] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0068] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0069] The present invention provides a remote training platform. Refer to Figure 1 , Figure 1 Schematic diagram of the structure of the operating environment of the remote training platform involved in the embodiment solution of the present invention.
[0070] As Figure 1 shown, the remote training platform includes: a processor 101, such as a CPU, a communication bus 102, a user interface 103, a network interface 104, and a memory 105. Among them, the communication bus 102 is used to realize the connection and communication between these components. The user interface 103 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). The network interface 104 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 105 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 105 may optionally also be a storage device independent of the aforementioned processor 101.
[0071] Those skilled in the art can understand that Figure 1 the hardware structure of the remote training platform shown in
[0072] As Figure 1 shown, the memory 105, as a computer-readable storage medium, may include an operating system, a network communication program module, a user interface program module, and a program for realizing the generation of remote training videos. Among them, the operating system schedules the communication between the various modules in the remote training platform and executes the program for generating remote training videos stored in the memory to realize the synthesis of training videos. Here, the synthesis includes the synthesis of the lecturer's portrait, teaching courseware, and the interactive screen between the lecturer and the students, which can greatly improve the real-time performance and interest of the training videos, and at the same time realize the on-site feeling of remote teaching.
[0073] In Figure 1In the hardware structure of the remote training platform shown, the network interface 104 is mainly used to access the network; the user interface 103 is mainly used to monitor the real-time teaching video of the instructor side and the real-time listening video of the student side. The videos of both sides are obtained through the monitoring of the user interface 103, and then the control processor 101 is used to call the generation program of the remote training video stored in the memory 105 to perform real-time synthesis on the monitored videos and update them to both sides, so as to enable both sides to observe the videos of each other in real time, and also reflect the integrity and on-site feeling of the teaching video, enhancing the interest of the teaching video. The specific implementation is as described in the operations of the embodiments of the remote training video generation method provided below.
[0074] Based on the above hardware structure of the remote training platform, various embodiments of the remote training video generation method of the present invention are proposed. Of course, the remote training platform listed here is only one implementation device for executing the remote training video generation method provided by the embodiments of the present invention. In practical applications, the implementation device can also be a training robot, which can be an AP or VR device. By executing this method, remote teaching can be realized, thereby enhancing the on-site experience of the training video screen, and at the same time, the training screen with two-way interaction can also be realized.
[0075] Refer to Figure 2 , Figure 2 is a flowchart of the remote training video generation method provided by the embodiments of the present invention. This method is essentially to realize the AI lecture hall for remote training. Specifically, by obtaining the teaching sample set of the instructor during the training process, the listening and interaction sample set of the students during the training process, and the teaching courseware sample set, a virtual AI lecture hall is constructed according to the teaching sample set, the listening and interaction sample set, the teaching courseware sample set, in combination with an image segmentation model and a video synthesis model.
[0076] During the training process, the specific personnel participating in the training are determined through face recognition, and then the video stream of the instructor during the teaching process is collected, the video stream of the students during the training process is collected, and the video stream of the courseware playback is collected. The video stream includes the teaching sample set of the instructor during the teaching process and the listening and interaction sample set of the students during the training process. The video stream is uploaded to the image segmentation model for image segmentation processing to obtain independent teaching samples of the instructor, independent listening and interaction samples of the students, and courseware samples. According to the independent teaching samples of the instructor, independent listening and interaction samples of the students, and courseware samples, a remote training AI lecture hall is generated by inputting them into the video synthesis model, thereby solving the geographical limitations of education, and at the same time, realizing the interactive experience between the lecturer and the standby lecturer. The remote training video generation method specifically includes the following steps:
[0077] Step S210, obtain the video stream during the training process, where the video stream includes: the teaching video when the instructor teaches and / or the listening and interaction video when the students participate in the training;
[0078] In this step, the video stream here should be understood as a general term for the video data generated from the instructor side and the student side respectively. The video stream may only include the video data of one end, or may include the video data of both ends. Specifically, it can be collected in real time through the Internet or read from the database of the transfer platform for remote training. When collecting from both ends respectively, first, a communication connection between the two ends needs to be established through the Internet or by logging in to a dedicated training account. This connection can be connected to the intermediate control platform or the two ends can be directly connected to each other.
[0079] If it is connected to the intermediate control platform, the method in this embodiment should be applied on the intermediate control platform. After finally synthesizing the video, the synthesized video is synchronously played and displayed on both ends in real time. If the two ends are directly connected to each other, the instructor side is selected as the main operation end, that is, the method provided in this embodiment is applied to the training device on the instructor side. The training device on the instructor side monitors the video stream on the student side, splits and extracts key information from the video stream on the student side and synthesizes it into the instructor's video stream, and synchronizes it to the student side for playback.
[0080] In practical applications, a depth camera is used to collect the video of the instructor's teaching during the remote training, extract the color and depth information of each frame in the video, and mark the area where the human figure is located to obtain the corresponding label. The convolutional neural network is trained through the color and depth information. Since there is color and spatial continuity of the human figure in the three-dimensional world, this convolutional neural network can be used to predict the area that may be the human figure in this video frame and segment it to obtain the video frame of the original picture of the teaching instructor.
[0081] Step S220, perform segmentation processing on the video stream through a preset image segmentation model to obtain independent pictures, where the independent pictures include the independent human figure picture of the instructor's teaching, the independent teaching courseware picture, and the independent student picture;
[0082] In practical applications, the image segmentation model here can specifically be the segmentation model of any one of the following algorithms: the threshold segmentation algorithm, the edge segmentation algorithm, the region segmentation algorithm, the image segmentation algorithm of cluster analysis, and the segmentation algorithm of artificial neural network. Based on these algorithms, training for image segmentation is carried out to obtain the final model.
[0083] In this embodiment, preferably, the segmentation algorithm of the artificial neural network is selected. Its basic idea is to obtain a linear decision function by training a multi-layer perceptron, and then use the decision function to classify pixels to achieve the purpose of segmentation. When using the artificial neural network method to segment images, a large amount of training data is required. There are a huge number of connections in the neural network, which is easy to introduce spatial information, thus further solving the problems of noise and non-uniformity in the image. In practical applications, as for which network structure to choose as the image segmentation model, it is specifically selected according to the actual needs of remote training.
[0084] Step S230: Extract the element content in the independent portrait screen, the independent teaching courseware screen, and / or the independent student screen, as well as the position information of the element content in the screen.
[0085] In this embodiment, when the acquired video stream is the video stream of the instructor side, the instructor teaching screen and the courseware screen in the video stream are segmented and separated. Specifically, the portrait recognition algorithm and the text recognition algorithm are used to identify and distinguish the portrait area and the courseware area. Based on this distinction, the video screens in the two areas are extracted. Of course, the portrait in the screen can also be extracted by the video matting technology, and the background color of the courseware area is filled to achieve the separation of the teaching screen and the courseware screen.
[0086] In practical applications, for the segmentation of these two types of screens, it can also be extracted by region annotation, that is, first identified, and then different identifiers are used to trace the edges of the region, and the movement of the region is monitored in real time. At the same time, the position information of the region is calculated, and this position information is for the covered area of the screen.
[0087] Step S240: Construct a picture frame of the virtual lecture hall according to the position information. The picture frame is a picture layout for simultaneously accommodating the independent portrait screen, the independent teaching courseware screen, and the independent student screen.
[0088] In this step, the picture frame here refers to a blank background picture, in which there are fixed wandering areas according to the position information of different element contents, or a connection relationship between the area and the portrait in the picture is established, and this area will move along with the portrait on the picture.
[0089] In practical applications, the area corresponding to the teaching courseware screen in the picture frame is fixedly set. Based on this fixed area, the movement relationship between the instructor portrait area and the student portrait area in the fixed area is established, and at the same time, the association between the movement area and the video screen is established to obtain the final picture frame.
[0090] Step S250, add the element content to the corresponding position of the screen frame to obtain a training video of the AI lecture hall.
[0091] In this step, when adding element content, it can also be automatically filled by presetting a fixed type of element to be filled in each position of the screen frame. That is, first establish the correspondence between positions and elements, and monitor the corresponding screen through the corresponding recognition technology to obtain the corresponding element content, and map the element content to the corresponding area in the screen frame, so that a training video screen showing any actions and information of both parties can be seen at the same time.
[0092] In practical applications, for the mapping of human figures, it can be only mapping their movement relationships, without mapping the actual human figure images. The mapped screen can be a preset small figure image. Of course, in order to improve interactivity and interest, it can be directly mapping real human figure images.
[0093] In this embodiment, there are two situations for the video stream of the lecturer side, that is, the teaching screen and the courseware screen can be recorded and obtained separately, or they can be recorded and obtained separately. If the lecturer uses projection for teaching, it is possible that the lecturer and the courseware are recorded and collected together. When the lecturer only uses a computer for teaching, generally the video screens of both are not recorded simultaneously, but separately.
[0094] In practical applications, the lecturer's teaching video is generally obtained in two parts, one part is the lecturer's lecture video, and the other part is the courseware video. And for these two situations, there are different IDs when recorded by the camera. Therefore, in order to facilitate video synthesis, in this embodiment, after obtaining the video stream, it is also necessary to detect the video stream. The specific process is as follows:
[0095] Detect whether the teaching video in the video stream is a mixed video, where the mixed video includes the lecturer's portrait video and the teaching courseware video;
[0096] If the teaching video is a mixed video, use the human face extraction algorithm to extract the independent portrait image of the lecturer in the face video, and use the character detection algorithm to extract the information of the courseware currently used by the lecturer in the teaching courseware video to obtain the independent teaching courseware image;
[0097] If the teaching video is not a mixed video, perform the step of segmenting the video stream through a preset image segmentation model to obtain independent images.
[0098] In practical applications, to detect whether the teaching video is a mixed video, it can specifically be determined by detecting whether the capture track of the teaching video is a single track. If so, it is determined that the teaching video is not a mixed video. Further, it can also be determined by detecting the video source in the teaching video. When it is detected that it does not belong to a mixed video, that is, the video in the video stream is captured from only one video source, it jumps to step S220 for video segmentation processing. Of course, before the segmentation processing, it can also be detected whether there is a portrait image in the video stream. If there is, step S220 is executed; otherwise, it directly jumps to step S230.
[0099] When it is detected that it belongs to a mixed video, that is, when there are video materials captured from two or more video sources in the video stream, the video stream is processed for image extraction through different extraction algorithms. The specific processing process is as Figure 3 shown.
[0100] For the acquisition of a mixed video stream, the specific method for extracting images is as follows:
[0101] S301, store the first source video and the second source video in the first video frame queue and the second video frame queue respectively in units of frames;
[0102] S302, extract an image of the first source video from the first video frame queue, process the image, extract the moving target of interest, and obtain the foreground image;
[0103] S303, extract an image of the second source video from the second video frame queue, and store the obtained foreground image and the image of the second source video extracted from the second video frame queue as different pictures respectively;
[0104] S304, repeat steps S302 - S303 until all the images in the first video frame queue and the second video frame queue are processed. Finally, synthesize the extracted images into a new picture set, and then synthesize the picture extracted from the student side into the picture set.
[0105] In this embodiment, in order to reduce the data processing of the remote training platform, specifically, when segmenting the teaching video, the video of the instructor side and the video of the student side can be processed separately. Preferably, they can be placed separately on the corresponding recording devices for processing. The specific processing for the instructor side is as follows:
[0106] When the video stream is the teaching video of the instructor, the steps of segmenting the video stream through a preset image segmentation model to obtain independent pictures include:
[0107] According to the preset depth - of - field formula, calculate the depth - of - field depth values of each picture in the teaching video;
[0108] Identify the foreground area and the background area of the picture in the teaching video according to the depth value, wherein the foreground area includes a portrait picture;
[0109] Use an image matting algorithm to extract the foreground area from the teaching video to obtain a foreground video picture, and extract the background area from the teaching video to obtain the independent teaching courseware picture;
[0110] According to a preset face recognition algorithm, identify the portrait picture in the foreground area, and extract the portrait picture from the foreground area to form the independent portrait picture.
[0111] In this embodiment, when the video stream is the student's listening and interaction video, the step of segmenting the video stream through a preset image segmentation model to obtain independent pictures includes:
[0112] Use face recognition technology to identify whether there are students in the listening and interaction video who meet the classroom interaction postures, where the classroom interaction postures include standing and raising hands;
[0113] If there is, perform a human body scan on the student through a camera to obtain the portrait contour of the student, and calculate the depth value of the portrait contour in the listening and interaction video according to a preset depth formula;
[0114] Use the depth value as the picture cutting critical point, and cut out all the pictures in the listening and interaction video located at the critical point to form the independent student picture.
[0115] In this embodiment, the obtained depth formula is as follows:
[0116] ,
[0117] ,
[0118] ,
[0119] where ΔL is the depth of field, δ is the allowable circle of confusion diameter, f is the lens focal length, F is the shooting aperture value of the lens, L is the focusing distance, and ΔL1 is the front depth of field.
[0120] In practical applications, the video streams generally collected from the trainee side are usually only the portrait videos of the trainees. Therefore, only the extraction of the portrait is required. Of course, it cannot be excluded that the images of the teaching courseware are also collected at the same time. For this situation, the face recognition technology is also used to extract the portrait images from the video stream, and the courseware does not need to be extracted. The teaching courseware images of the instructor side can be used when synthesizing the video, which can reduce the error probability of the information in the synthesized video.
[0121] In this embodiment, when there is interaction with the trainees, it is also necessary to record the interaction information of the trainees, that is to say, it is also necessary to extract various action information and question information of the trainees through action tracking.
[0122] After calculating the respective depths of field in the video stream based on the above calculation formula, the independent portrait images of the instructor's teaching, the independent teaching courseware images, and the independent trainee images are extracted based on the depths of field. In practical applications, the independent portrait images and the independent teaching courseware images are generally obtained from the video stream of the instructor side. That is to say, the depth of field where the portrait of the instructor is located is the foreground area, and the portrait image of the instructor is obtained from the foreground area, while the depth of field value with a relatively larger value is the background area, and the courseware information is extracted from it.
[0123] In order to improve the quality of the training video, in this embodiment, after the corresponding images are extracted, important information on teaching and listening is also extracted from the corresponding images through matte extraction and text extraction technologies, and a new video image is constructed based on the important information, that is, the element content in the independent portrait images, the independent teaching courseware images, and / or the independent trainee images, as well as the position information of the element content in the image.
[0124] In practical applications, the specific implementation process of this extraction step is as Figure 4 shown:
[0125] Step S401, create a canvas according to the length and width of the independent image, and select any corner point in the canvas as the coordinate origin to establish a two-dimensional coordinate system;
[0126] Step S402, based on the two-dimensional coordinate system, calculate the coordinate information of the portrait of the instructor or the portrait of the trainee in the independent image, and calculate the coordinate information of the teaching content of the teaching courseware in the independent image;
[0127] Step S403, extract the portrait and the courseware content from the independent image according to the coordinate information.
[0128] In practical applications, when creating a two-dimensional coordinate system, it is usually constructed based on the display device of the image, and for those with different display sizes, conversion can be performed according to the resolution to obtain the corresponding position information.
[0129] In this embodiment, for the extraction of elements, it can also be distinguished and extracted according to color. For example, the background color is recognized, and all elements in the background color are extracted as courseware information, and the elements that are not the background color are extracted as portrait elements. Of course, for the extraction of portrait elements, it can also be determined whether they are moving. If they are moving, they are considered portrait elements, and then the outline is depicted and extracted.
[0130] Furthermore, before constructing a new training video screen based on the extracted elements, it is also necessary to construct the screen layout of the classroom. The specific construction steps are as Figure 5 shown:
[0131] Step S501: Use the independent teaching courseware screen as the background canvas of the AI lecture hall, and construct a coordinate system on the background canvas;
[0132] Step S502: According to the coordinate information, draw a portrait filling area on the background canvas that is the same shape as the portrait to obtain the screen frame;
[0133] Based on the constructed screen frame, the extracted elements can be filled into the corresponding positions. Specifically, the steps of adding the element content to the corresponding positions of the screen frame to obtain the training video of the AI lecture hall include:
[0134] Fill the extracted portrait into the corresponding portrait filling area, and fuse the portrait with the background canvas through the boundary interpolation background fusion algorithm to obtain the training video.
[0135] In practical applications, in the process of synthesizing the extracted elements into a training video, a video synthesis model can be used. Specifically, after the video synthesis model synthesizes the elements in the screen frame into a complete video, it also includes retouching the filled training video, that is, fusing the edges in the screen frame through noise reduction, feathering, etc., so as to achieve seamless connection of video materials.
[0136] In summary, by using the method provided in this embodiment to generate a training video, the participation of students in the remote training scenario and the interaction effect between the lecturer and students can be improved, and finally assist schools, institutions, etc. to improve the training effect and students' academic performance in remote training, teaching and other links.
[0137] Next, taking the teaching video of the lecturer side as an example, the implementation of the training video generation method provided by the present invention will be described in detail, as Figure 6 shown.
[0138] Step S601: Use a depth camera to collect the video of the lecturer teaching during the remote training;
[0139] Among them, the video here includes the portrait of the lecturer and the courseware information. Preferably, the courseware information here can be directly read from the training device. Of course, it can also be obtained by recording from the projection screen of the training video using a depth camera.
[0140] Step S602: Extract the color and depth information of each frame in the video, and label the area where the portrait is located to obtain the corresponding label.
[0141] In this embodiment, when extracting video frames, specifically, a convolutional neural network can be trained through color and depth information. Based on the trained neural network, the color and depth information in the video is extracted, and the continuity in the three-dimensional space among the color, depth information, and the label is also established. Further, the continuity in space with the lecturer's portrait can also be added. This convolutional neural network can be used to predict the area in the video frame that may be the portrait, and segment it to obtain the video frame of the original picture of the lecturer during the lecture.
[0142] Step S603: Superimpose the obtained lecturer during the lecture and the set background image, and perform interpolation denoising on the boundary to obtain the generated video frame.
[0143] Step S604: Combine according to the time sequence of the video frames to obtain the AI lecture hall.
[0144] The invention of the remote training AI lecture hall based on image segmentation and video synthesis extracts the lecturer's teaching video and the behavior videos such as the interaction of students through image segmentation, and then combines the courseware video through the method of real-time online video generation to generate the AI lecture hall, which can greatly improve the participation of students in the remote training scenario and the interaction effect between the lecturer and the students, and finally assist schools, institutions, etc. to improve the training effect and the learning performance of students in the links of remote training, teaching, etc.
[0145] In this embodiment, if there is also the interaction video of the students at the same time, it can also add the real-time video information of the students' listening to the lecture to the synthesized AI classroom video. The implementation process is as follows:
[0146] First, during the training process, determine the specific personnel participating in the training through face recognition, collect the video stream of the lecturer during the lecture, collect the video stream of the students during the training, and collect the video stream of the courseware playback. The video stream includes the teaching sample set of the lecturer during the lecture and the listening and interaction sample set of the students during the training.
[0147] Then, upload the video stream to the image segmentation model for image segmentation processing to obtain the independent sample of the lecturer's teaching, the independent sample of the students' listening and interaction, and the courseware sample.
[0148] Generate a remote training AI classroom according to the instructor teaching independent sample, the student listening interaction independent sample, and the courseware sample input video synthesis model.
[0149] At this time, during the process of the video synthesis model synthesizing the video, it is also necessary to compare and deduplicate the instructor teaching independent sample, the student listening interaction independent sample, the courseware sample with the instructor teaching independent sample and the courseware sample received from the instructor side, and synthesize the different video frames into the video, so as to ensure the conciseness and accuracy of the video, or the real-time nature of the video.
[0150] In summary, the remote training AI classroom based on image segmentation and video synthesis of the present invention uses an image segmentation model and a video synthesis model to analyze and process the process of students participating in training and the process of instructors teaching, and generates a remote training AI classroom, which can enhance the sense of participation of students in the remote training scenario and the interaction effect between instructors and students, and ultimately assist schools, institutions, etc. to improve the training effect in the links of remote training and teaching.
[0151] To solve the above problems, an embodiment of the present invention also provides a device for generating a remote training video, as Figure 7 shown, the device for generating a remote training video includes:
[0152] An acquisition module 71, configured to obtain a video stream during the training process, where the video stream includes: a teaching video when the instructor teaches and / or a listening interaction video when the student participates in training;
[0153] A segmentation module 72, configured to perform segmentation processing on the video stream through a preset image segmentation model to obtain independent pictures, where the independent pictures include an independent portrait picture of the instructor teaching, an independent teaching courseware picture, and an independent student picture;
[0154] An extraction module 73, configured to extract the element content in the independent portrait picture, the independent teaching courseware picture and / or the independent student picture, and the position information of the element content in the picture;
[0155] A synthesis module 74, configured to construct a picture frame of a virtual classroom according to the position information, where the picture frame is a picture layout for simultaneously accommodating the independent teaching picture, the independent teaching courseware picture, and the independent student picture; add the element content to the corresponding position of the picture frame to obtain a training video of the AI classroom.
[0156] Based on the execution function of this device and the execution process corresponding to the function, it is the same as the content described in the embodiment of the method for generating a remote training video of the above embodiment of the present invention. Therefore, this embodiment does not elaborate too much on the content of the embodiment of the device for generating a remote training video.
[0157] In addition, an embodiment of the present invention further provides a training device, which includes: a memory, a processor, and a generation program of a remote training video stored on the memory and executable on the processor. The method implemented when the generation program of the remote training video is executed by the processor may refer to various embodiments of the method for generating a remote training video of the present invention, and thus will not be elaborated herein.
[0158] The present invention also provides a computer-readable storage medium.
[0159] In this embodiment, a generation program of a remote training video is stored on the computer-readable storage medium. The method implemented when the generation program of the remote training video is executed by a processor may refer to various embodiments of the method for generating a remote training video of the present invention, and thus will not be elaborated herein.
[0160] In the method and device provided by the embodiments of the present invention, an image segmentation model and a video synthesis model are mainly used to analyze and process the training process of trainees and the teaching process of instructors, generate a remote training AI classroom, so as to enhance the sense of participation of trainees in the remote training scenario and the interaction effect between instructors and trainees.
[0161] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM) and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0162] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose and scope protected by the claims of the present invention. All those that adopt equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, all fall within the protection scope of the present invention.
Claims
1. A method for generating remote training videos, applied to a remote training platform, characterized in that, The method for generating the remote training video includes the following steps: Obtain the video stream during the training process, where the video stream includes: the teaching video when the instructor is teaching and / or the interactive video of the students attending the training; Perform segmentation processing on the video stream through a preset image segmentation model to obtain independent pictures, where the independent pictures include one or more of the independent portrait pictures of the instructor teaching, the independent teaching courseware pictures, and the independent student pictures; Among them, the step of performing segmentation processing on the video stream through a preset image segmentation model to obtain independent pictures includes: When the video stream is the teaching video of the instructor, calculate the depth values of the depths of the pictures in the teaching video according to a preset depth of field formula; Identify the foreground area and the background area of the pictures in the teaching video according to the depth values, where the portrait picture is included in the foreground area; Use an image matting algorithm to extract the foreground area from the teaching video to obtain a foreground video picture, and extract the background area from the teaching video to obtain the independent teaching courseware picture; Identify the portrait picture in the foreground area according to a preset portrait recognition algorithm, and extract the portrait picture from the foreground area to form the independent portrait picture; Extract the element content in the independent picture and the position information of the element content in the picture, where it includes creating a canvas according to the length and width of the independent picture, and selecting any corner point in the canvas as the coordinate origin to establish a two-dimensional coordinate system; based on the two-dimensional coordinate system, calculate the coordinate information of the portrait of the instructor or the portrait of the student in the independent picture, and calculate the coordinate information of the teaching content of the teaching courseware in the independent picture; according to the coordinate information, extract the portrait and the courseware content from the independent picture; Construct a picture frame of a virtual lecture hall according to the position information, where the picture frame is a picture layout for accommodating the independent pictures; Add the element content to the corresponding positions of the picture frame to obtain the training video of the AI lecture hall.
2. The method for generating a remote training video according to claim 1, wherein After the step of obtaining the video stream during the training process, it further includes: Detect whether the teaching video in the video stream is a mixed video, where the mixed video includes the portrait video of the instructor and the teaching courseware video; If the teaching video is a mixed video, use a portrait extraction algorithm to extract the independent portrait picture of the instructor in the portrait video, and use a character detection algorithm to extract the courseware information currently used by the instructor in the teaching courseware video, and synthesize the independent portrait picture and the courseware information into an independent teaching courseware picture; If the teaching video is a non-mixed video, perform the step of performing segmentation processing on the video stream through a preset image segmentation model to obtain independent pictures.
3. The method for generating a remote training video according to claim 1, wherein, The step of performing segmentation processing on the video stream through a preset image segmentation model to obtain independent pictures includes: When the video stream is the video of the student's classroom interaction, the face recognition technology is used to identify whether there is a student in the classroom interaction video who meets the classroom interaction postures, where the classroom interaction postures include standing and raising hands; If there is, the camera performs a human body scanning process on the student to obtain the portrait contour of the student, and calculates the depth of field value of the portrait contour in the classroom interaction video according to a preset depth of field formula; Taking the depth of field value as the picture cutting critical point, all the pictures located on the critical point in the classroom interaction video are cut out to form the independent student picture.
4. The method for generating a remote training video according to claim 1, wherein The step of constructing the picture frame of the virtual lecture hall according to the position information includes: Using the independent teaching courseware picture as the background canvas of the AI lecture hall, and constructing a coordinate system on the background canvas; According to the coordinate information, a portrait filling area with the same shape as the portrait is outlined on the background canvas to obtain the picture frame; The step of adding the element content to the corresponding position of the picture frame to obtain the training video of the AI lecture hall includes: Filling the extracted portrait into the corresponding portrait filling area, and fusing the portrait with the background canvas through the boundary interpolation background fusion algorithm to obtain the training video.
5. The method for generating a remote training video according to claim 4, wherein, The depth of field calculation formula is: , Where δ is the allowable circle of confusion diameter, f is the lens focal length, F is the shooting aperture value of the lens, and L is the focusing distance.
6. A generating device for remote training videos, characterized in that, The generating device of the remote training video includes: An acquisition module, configured to acquire a video stream during the training process, where the video stream includes: the teaching video when the lecturer teaches and / or the classroom interaction video when the student participates in the training; A segmentation module, configured to perform segmentation processing on the video stream through a preset image segmentation model to obtain independent pictures, where the independent pictures include one or more of the independent portrait picture of the lecturer teaching, the independent teaching courseware picture, and the independent student picture; where the step of performing segmentation processing on the video stream through a preset image segmentation model to obtain independent pictures includes: when the video stream is the teaching video of the lecturer, calculating the depth value of the depth of field of each picture in the teaching video according to a preset depth of field formula; identifying the foreground area and the background area of the pictures in the teaching video according to the depth value, where the foreground area includes the portrait picture; using an image matting algorithm to extract the foreground area from the teaching video to obtain a foreground video picture, and extracting the background area from the teaching video to obtain the independent teaching courseware picture; identifying the portrait picture in the foreground area according to a preset portrait recognition algorithm, and extracting the portrait picture from the foreground area to form the independent portrait picture; An extraction module, configured to extract the element content in the independent picture and the position information of the element content in the picture, including creating a canvas according to the length and width of the independent picture, selecting any corner point in the canvas as the coordinate origin, and establishing a two-dimensional coordinate system; based on the two-dimensional coordinate system, calculating the coordinate information of the portrait of the lecturer or the portrait of the student in the independent picture, and calculating the coordinate information of the teaching content of the teaching courseware in the independent picture; according to the coordinate information, extracting the portrait and the courseware content from the independent picture. A synthesis module, configured to construct a picture framework of a virtual lecture hall according to the position information, where the picture framework is a picture layout for accommodating the independent picture; adding the element content to the corresponding position of the picture framework to obtain a training video of an AI lecture hall.
7. A training device, characterized in that, The training device includes: a memory, a processor, and a generation program of a remote training video stored on the memory and executable on the processor. When the generation program of the remote training video is executed by the processor, the steps of the method for generating a remote training video according to any one of claims 1-5 are implemented.
8. A computer-readable storage medium, characterized in that, A generation program of a remote training video is stored on the computer-readable storage medium. When the generation program of the remote training video is executed by a processor, the steps of the method for generating a remote training video according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Multimedia interaction teaching system and teaching method
CN104469089A
Image segmentation method and mobile terminal
CN107507239A
Video processing method and device and storage medium
CN110290425A