An image processing method, device, electronic equipment and readable storage medium
Patent Information
- Application Number
- CN202310363832.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-06
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-04-06
AI Technical Summary
[0005]本发明实施例提供一种图像处理方法、装置、电子设备及可读存储介质,以解决现有技术中通过对图像进行内容识别与捕捉进行导播处理时,由于分析计算量过大造成导播处理的应用效果较差的问题
[0020]本发明实施例中,首先获取原始的视频流;根据预设的提取规则对视频流的图像帧进行兴趣区域提取,获得多个兴趣区域;之后将从图像帧中提取的多个兴趣区域的内容,输入至内容处理模型,获得针对每个兴趣区域的导播权重值;并在播放视频流的过程中,将导播权重值满足预设条件的目标兴趣区域内的图像内容,进行导播处理。对于整个视频流的导播处理,不再完整计算整个视频画面的所有内容,而是优先根据导播权重值划分出多个不同的兴趣区域,只针对兴趣区域内的部分关键内容进行分析计算等操作,省去了对于非兴趣区域的计算过程,减轻了图像数据的计算量,使得导播处理的应用效果更佳。
Smart Images

Figure CN116527828B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia technology, and in particular to an image processing method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] As a digital video media technology, broadcasting has important applications in television programs and situational variety show production. Now, in the education sector, with the continuous improvement of smart device infrastructure, broadcasting technology is also gradually demonstrating its technological advantages in applications such as live streaming and recorded courses in smart classrooms.
[0003] In existing technologies, smart classroom broadcasting systems based on artificial intelligence algorithms can identify and capture areas of scene change in the classroom footage captured by deep learning models, and provide close-ups of these areas through broadcasting, recording exciting content from a more focused perspective.
[0004] However, in the existing solutions, the relevant algorithms need to analyze and calculate all images in the monitoring screen during the content recognition and capture process. This results in a large amount of analysis and calculation, leading to poor application effects for broadcasting. Summary of the Invention
[0005] This invention provides an image processing method, apparatus, electronic device, and readable storage medium to solve the problem that the application effect of broadcast processing is poor due to the excessive amount of analysis and calculation when performing content recognition and capture of images in the prior art.
[0006] In a first aspect, embodiments of the present invention provide an image processing method, the method comprising:
[0007] Obtain the raw video stream;
[0008] According to preset extraction rules, the image frames of the video stream are subjected to region of interest extraction to obtain multiple regions of interest;
[0009] The content of multiple regions of interest extracted from the image frame is input into the content processing model to obtain the director weight value for each region of interest;
[0010] During the playback of the video stream, the image content within the target interest region whose director weight value meets the preset conditions is subjected to director processing.
[0011] Secondly, embodiments of the present invention provide an image processing apparatus, characterized in that the apparatus comprises:
[0012] The video stream acquisition module is used to acquire the raw video stream;
[0013] The region of interest (ROI) acquisition module is used to extract ROIs from the image frames of the video stream according to preset extraction rules, thereby obtaining multiple ROIs.
[0014] The director weight value determination module is used to input the content of multiple interest regions extracted from the image frame into the content processing model to obtain the director weight value for each interest region;
[0015] The directing processing execution module is used to perform directing processing on the image content within the target interest area whose directing weight value meets preset conditions during the playback of the video stream.
[0016] Thirdly, embodiments of the present invention provide an electronic device, including: a processor;
[0017] Memory used to store the processor's executable instructions;
[0018] The processor is configured to execute the instructions to implement the method.
[0019] Fourthly, embodiments of the present invention provide a readable storage medium that, when instructions in the readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the method.
[0020] In this embodiment of the invention, the original video stream is first acquired; then, regions of interest (ROIs) are extracted from the image frames of the video stream according to preset extraction rules to obtain multiple ROIs; subsequently, the content of the multiple ROIs extracted from the image frames is input into a content processing model to obtain a director weight value for each ROI; and during the playback of the video stream, the image content within the target ROI region whose director weight value meets preset conditions is subjected to director processing. For the director processing of the entire video stream, the entire content of the video frame is no longer calculated completely. Instead, multiple different ROIs are first divided according to the director weight values, and only key content within the ROI regions is analyzed and calculated. This eliminates the calculation process for non-ROI regions, reduces the computational load of image data, and makes the application effect of director processing better.
[0021] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0022] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0023] Figure 1 This is a simplified flowchart of the implementation steps of an image processing method provided in an embodiment of the present invention;
[0024] Figure 2 This is a schematic diagram of a teacher-side video broadcast screen provided in an embodiment of the present invention;
[0025] Figure 3 This is a schematic diagram of a student-side video broadcast screen provided in an embodiment of the present invention;
[0026] Figure 4 This is a complete flowchart of the implementation steps of an image processing method provided in an embodiment of the present invention;
[0027] Figure 5 This is a dynamic schematic diagram of region of interest extraction provided by an embodiment of the present invention;
[0028] Figure 6 This is a schematic diagram of a region image size adjustment process provided in an embodiment of the present invention;
[0029] Figure 7 This is an execution logic block diagram of an image processing method provided in an embodiment of the present invention;
[0030] Figure 8 This is a diagram illustrating the effect of emphasizing realism in an embodiment of the present invention;
[0031] Figure 9 This is another implementation effect diagram provided by the embodiment of the present invention, emphasizing the realistic effect;
[0032] Figure 10 This is another implementation effect diagram provided by the embodiment of the present invention, emphasizing the realistic effect;
[0033] Figure 11 This is a schematic diagram of the functional components of an image processing device provided in an embodiment of the present invention;
[0034] Figure 12 This is a functional component relationship diagram of an electronic device provided in an embodiment of the present invention;
[0035] Figure 13 This is a functional component relationship diagram of another electronic device provided in an embodiment of the present invention. Detailed Implementation
[0036] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are illustrated in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the invention and to fully convey the scope of the invention to those skilled in the art.
[0037] Reference Figure 1 This diagram illustrates a simplified implementation flow of an image processing method provided by an embodiment of the present invention. Figure 1 As shown, the steps of the method include:
[0038] Step 101: Obtain the raw video stream.
[0039] The image processing method provided in this invention is used for video broadcasting processing, including live streaming and recorded broadcasting. First, it is necessary to acquire the video material to be processed, i.e., the original video stream. In specific implementation, since the scene frame that a single camera can record is limited, and it is impossible to simultaneously capture images from multiple different angles in a short period of time, the video stream to be processed generally comes from multiple camera devices.
[0040] In the image processing method provided in this embodiment of the invention, it is mainly applied to a smart classroom broadcasting system based on Artificial Intelligence (AI), which requires accurate judgment of every student and teacher in the classroom. Therefore, the broadcasting camera needs to switch between a panoramic view of the classroom and a Region of Interest (ROI) with content value, thereby achieving fixed-position broadcasting processing. During this process, the positions where close-up shots can be taken are certain fixed ROI regions defined according to the actual scene.
[0041] Step 102: Extract regions of interest from the image frames of the video stream according to preset extraction rules to obtain multiple regions of interest.
[0042] The acquired video stream image data is read and analyzed, and different regions are divided according to the ROI.
[0043] For example, in an embodiment of the present invention, for a smart classroom equipped with a broadcast control system, the system includes acquiring video streams for students and video streams for teachers. (See reference...) Figure 2 This illustration shows a schematic diagram of a teacher-side video broadcast screen provided by an embodiment of the present invention; as shown below. Figure 2As shown, the current view includes four generated areas of interest: the display screen area 201 excluding the podium, the first blackboard area 202, the second blackboard area 203, and the teacher's area 204. Among them, the display screen area 201 is slightly larger than the display screen 211.
[0044] Same reference Figure 3 This illustration shows a schematic diagram of a student-side video broadcast screen provided by an embodiment of the present invention; as shown below. Figure 3 As shown, the central area of the current screen view is the student seating area, with aisles on both sides. The central part includes multiple student areas, namely the second student area 302, the third student area 303, the fifth student area 305, the sixth student area 306, and the seventh student area 307. In addition, there are the first student area 301 and the fourth student area 304 on the sides of the screen.
[0045] It is worth noting that there may be overlapping areas among multiple regions of interest. Figure 2 and Figure 3 The region of interest (ROI) division shown is defined by the developers based on the specific scenario of a smart classroom. However, in the image processing method provided in this embodiment of the invention, when performing image analysis on the acquired video stream, the acquired ROI can be adjusted according to different broadcasting scenarios; this embodiment of the invention does not impose any limitations here.
[0046] Step 103: Input the content of multiple regions of interest extracted from the image frame into the content processing model to obtain the director weight value for each region of interest.
[0047] After extracting multiple regions of interest (ROIs) from the video stream, the image content within each ROI is input into a content processing model for evaluation, resulting in a director's weight value for each ROI. This director's weight value is used to assess the priority of events occurring within the ROI.
[0048] Since the purpose of a director's close-up is to give a close-up of a certain area of the panoramic view of the entire video stream, how to select the area of interest worth paying attention to from multiple areas of interest needs to be considered according to different scene characteristics.
[0049] Specifically, during a teaching session, the focus of classroom activity is primarily on the teacher. The focus is on the teaching content: the teacher's explanations at the podium, the handwritten notes on the blackboard, and the slides displayed on the screen are all part of the teaching content. Since the pace of classroom activities is mainly controlled by the teacher, only one of these scenarios needs to be monitored at any given time.
[0050] Furthermore, if the content displayed on the blackboard and slides remains unchanged for a certain period, and the teaching content is mainly conveyed orally by the teacher, then the area where the teacher is located will have the highest attention value, and therefore the guiding characteristic value of this area of interest will be high. Subsequently, if new content is added to the blackboard or the slides change, the attention value of the areas where these events occur should increase. Additionally, from the student's perspective, if a student raises their hand to answer a question during the lesson, the guiding priority for the area where that student is located will also be high.
[0051] Step 104: During the playback of the video stream, the image content within the target interest region whose director weight value meets the preset conditions is subjected to director processing.
[0052] Among multiple regions of interest, a target region of interest that meets preset conditions is determined based on the obtained director weight values, and then director processing is performed. In this embodiment of the invention, a corresponding weight threshold is preset for the director weight values. If the director weight value obtained for a determined region of interest exceeds the weight threshold, the region of interest is determined as the target region of interest.
[0053] Specifically, for example, if the weight threshold is set to 65, and the teacher's area 204 has a director weight of 70, the display screen area 201 has a weight of 60, and the first blackboard area 202 and the second blackboard area 203 have weights of 55, then the teacher's area 204 in the current video stream is determined to be the target area of interest.
[0054] There are several specific methods for implementing broadcasting processing for a target area of interest. For example, the screen of the target area of interest can be magnified so that the entire display device only plays the content of the target area of interest during the viewing process.
[0055] It's worth noting that directing isn't limited to continuously selecting a specific area of interest during video playback. Instead, it involves switching between multiple areas of interest and panoramic views of the entire scene, using camera language to guide viewers to focus on key content based on the events occurring in each area and the overall rhythm of the scene.
[0056] In summary, the image processing method provided by this invention first acquires the original video stream; then, it extracts regions of interest (ROIs) from the image frames of the video stream according to preset extraction rules, obtaining multiple ROIs; next, it inputs the content of the multiple ROIs extracted from the image frames into a content processing model to obtain a director weight value for each ROI; and during the playback of the video stream, it performs director processing on the image content within the target ROI region whose director weight value meets preset conditions. For the director processing of the entire video stream, instead of calculating all the content of the entire video frame, it prioritizes dividing multiple different ROIs according to the director weight values, and only performs analysis and calculation on key content within the ROI regions, eliminating the calculation process for non-ROI regions, reducing the computational load of image data, and making the application effect of director processing better.
[0057] Reference Figure 4 This diagram illustrates a complete implementation flow of an image processing method provided by an embodiment of the present invention; as follows: Figure 4 As shown, the steps of the method include:
[0058] Step 401: Obtain the raw video stream.
[0059] For details of this step, please refer to step 101 above. This embodiment will not repeat the details here.
[0060] Optionally, in one embodiment, step 401 may specifically include:
[0061] Sub-step 4011: Obtain the student-side video stream from the perspective of the student's location, and the teacher-side video stream from the perspective of the teacher's location.
[0062] Reference Figure 2 and Figure 3 The images show the video stream displayed from the teacher's location and the video stream displayed from the student's location.
[0063] It should be noted that the video stream described above is only the video material obtained for the application scenario of a smart classroom provided in this embodiment of the invention. In practical applications, there are also video streams from various other perspectives, which are not limited here.
[0064] Step 402: Extract regions of interest from the image frames of the video stream according to preset extraction rules to obtain multiple regions of interest.
[0065] For details of this step, please refer to step 102 above. This embodiment will not repeat the details here.
[0066] Optionally, in one embodiment, step 402 may specifically include:
[0067] Sub-step 4021: Perform region extraction on the student-end video stream to obtain multiple regions of interest including image frames of the student-end video stream, and perform region extraction on the teacher-end video stream to obtain multiple regions of interest including image frames of the teacher-end video stream.
[0068] like Figure 2 As shown, the current view includes four generated areas of interest: the display screen area 201 excluding the podium, the first blackboard area 202, the second blackboard area 203, and the teacher's area 204. Among them, the display screen area 201 is slightly larger than the display screen 211.
[0069] Similarly, Figure 3 As shown, the central area of the current screen view is the student seating area, with aisles on both sides. The central part includes multiple student areas, namely the second student area 302, the third student area 303, the fifth student area 305, the sixth student area 306, and the seventh student area 307. In addition, there are the first student area 301 and the fourth student area 304 on the sides of the screen.
[0070] In an optional embodiment, sub-step 4021 may further include:
[0071] Sub-step 40211: Perform object recognition on the image frame to obtain an object recognition box.
[0072] The extraction process for regions of interest in this embodiment of the invention primarily follows the extraction rules preset by the developers. In smart classroom application scenarios, the panoramic view of the teacher's and student's ends is generally a fixed viewpoint, and the types of events requiring focused display are limited, for example... Figure 2 The content in the four interest areas shown is sufficient to meet the directing requirements of this scenario using preset fixed areas.
[0073] like Figure 2 As shown, the current viewpoint includes four generated regions of interest: the display screen area 201, the first blackboard area 202, the second blackboard area 203, and the teacher's area 204. The display screen and blackboard have fixed geometric dimensions and fixed relative positions within the teacher's viewpoint. Object recognition boxes can be directly obtained by identifying the geometric edges of these objects, and these boxes will serve as the edge lines of the regions of interest. Furthermore, for the classroom area 204, representing the teacher's location, due to the uncertainty of personnel movement, using a fixed area lacks versatility. Therefore, it is necessary to perform object recognition on the personnel and use the obtained object recognition boxes as the edge lines of the regions of interest.
[0074] Sub-step 40212: In response to the selection operation of the object recognition box, select the area where part of the object recognition box is located as the region of interest.
[0075] After obtaining object bounding boxes through object recognition of the video stream, the region containing at least a portion of the object bounding boxes is selected as the region of interest based on the actual application scenario. For example, refer to... Figure 5 The diagram illustrates a dynamic illustration of region of interest extraction according to an embodiment of the present invention. Six students in two adjacent rows of seats generate object recognition boxes through object recognition. Based on these, six object recognition boxes within this range are selected to form a region of interest, as shown by the dashed box. The region of interest completely encompasses the generated object recognition boxes.
[0076] Optionally, in one embodiment, the number of regions of interest is directly proportional to the density of the object recognition boxes. Specifically, for students, the higher the student density within the classroom, the more regions of interest will be extracted. During this process, even if some regions of interest overlap significantly, no two regions will have exactly the same size and location.
[0077] Sub-step 4022: Extract the region of interest (ROI) of the first image frame in the video stream according to the preset extraction rules, and obtain the ROI of the first image frame.
[0078] For video stream directing, it is necessary to analyze all content in the entire video stream. Similarly, the regions of interest identified in the above steps also need to be extracted from the entire video stream.
[0079] Specifically, in this embodiment of the invention, since the video stream is formed by playing multiple static image frames in chronological order, the region of interest (ROI) of the first image frame in the video stream is first extracted according to a preset extraction rule to obtain the ROI of the first image frame. For specific results, refer to... Figure 2 and Figure 3 The details of this embodiment will not be elaborated here.
[0080] Sub-step 4023: Map the contour position of the region of interest in the image frame at the first moment to the image frames at all other moments in the video stream to obtain the region of interest for each image frame.
[0081] After obtaining the region of interest for the first frame, since the panoramic image size of the entire video stream is the same, the outline position of the region of interest in the first frame can be directly mapped to the image frames of all other frames in the video stream. The size and position of the mapped region of interest are exactly the same as the region of interest in the first frame. In this way, the region of interest for each image frame can be obtained.
[0082] Step 403: Adjust the size of each of the multiple regions of interest, thereby adjusting the screen size of each of the multiple regions of interest to the same preset screen size.
[0083] Reference Figure 6 The diagram illustrates a regional image size adjustment process provided by an embodiment of the present invention. Figure 6 As shown, the current video stream contains two extracted regions of interest: Region of Interest-1 and Region of Interest-2. Before inputting the content of these regions of interest into the content processing model, the system also needs to adjust the size of the images within each region of interest.
[0084] Specifically, in step 402, when extracting regions of interest using preset extraction rules, the size of all regions of interest is not strictly limited to the same size for the sake of flexibility in region extraction. However, when further processing the regions of interest, the processing model algorithm used has uniformity. Adjusting the screen size of all regions of interest to the same size facilitates parallel computation in subsequent content processing and reduces unnecessary resource waste.
[0085] Reference Figure 6 After resizing the image content of the region of interest, the images of individual regions of interest are combined side-by-side into an image packet. Thanks to the parallel design of the graphics processing unit (GPU), using parallel data packets can reduce the number of layers to be processed and reduce computation time.
[0086] Step 404: Input the content of the multiple regions of interest extracted from the image frame into the content processing model to obtain the director weight value for each region of interest.
[0087] For details of this step, please refer to step 103 above. This embodiment will not repeat the details here.
[0088] Optionally, in one embodiment, the broadcast weight value of the region of interest is proportional to the broadcast priority of the image content within the region of interest.
[0089] Step 405: During the playback of the video stream, the image content within the target interest area whose director weight value meets the preset conditions is subjected to director processing.
[0090] For details of this step, please refer to step 104 above. This embodiment will not repeat the details here.
[0091] Optionally, in one embodiment, step 405 may specifically include:
[0092] Sub-step 4051: Sort the regions of interest according to the director weight value and obtain the region of interest score sequence.
[0093] After obtaining the director weight values for multiple regions of interest (ROIs) through the content processing model, they are simply sorted according to their numerical values to obtain a ROI score sequence. Generally, for the same video stream at any given time, only one ROI can be processed using director methods; therefore, it is necessary to filter target ROIs that meet preset conditions based on their director weight values.
[0094] Sub-step 4052: In the interest region score sequence, if it is determined that the director weight value of the interest region is greater than the preset weight threshold, the interest region whose director weight value is greater than the weight threshold is determined as the target interest region.
[0095] In an optional embodiment, if it is determined that in the region of interest score sequence, the director weight values of multiple regions of interest are all greater than the weight threshold, the method further includes:
[0096] Sub-step 40521: Obtain the content event type of the image frame within the region of interest, and the event priority corresponding to the content event type; the content event type is used to characterize the features of events occurring within the target region of interest.
[0097] In this embodiment of the invention, any region of interest whose obtained director weight value exceeds a preset weight threshold is considered qualified for directorial work. However, in practical applications, it is inevitable that multiple regions of interest may simultaneously have director weight values exceeding the preset weight threshold, thus requiring further selection.
[0098] Different content events occur in different interest areas, and different event priorities can be set for different content events. For example, in a student scenario, the priority order can be set as follows: standing up to answer a question -> raising a hand -> listening attentively -> looking down and fidgeting. In this way, it is possible to select the more noteworthy content events from multiple interest areas that have the authority to handle the broadcast.
[0099] Sub-step 40522: Among all the interest regions whose director weight value is greater than the weight threshold, the interest region with the highest priority of the event type is determined as the target interest region, and the image content within the target interest region is subjected to director processing.
[0100] Following sub-step 40521, among all regions of interest whose director weight value is greater than the weight threshold, the region of interest with the highest event type priority is determined as the target region of interest, and the image content within the target region of interest is processed by the director.
[0101] Sub-step 4053: Perform broadcasting processing on the image content within the target area of interest.
[0102] Reference Figure 7 This diagram illustrates the execution logic block diagram of an image processing method provided by an embodiment of the present invention. First, a raw video stream is acquired using a video capture device (generally a camera device). Then, regions of interest (ROIs) are extracted according to preset extraction rules. After adjusting the size of the images of the multiple ROIs, the content of the ROIs is evaluated using a content processing model to generate director weight values. Based on the director weight values obtained for each ROI, the ROI that conforms to the preset rules is selected as the target ROI for close-up output.
[0103] There are several different methods for implementing broadcast control.
[0104] Alternatively, in one embodiment, sub-step 4053 may further include:
[0105] Sub-step 40531: Use the geometric center of the target region of interest as the target magnification center; during the playback of the video stream, use the target magnification center as a reference to uniformly magnify the image content in the target region of interest to a preset size for display within a first duration.
[0106] Reference Figure 8 The diagram illustrates the effect of a content-emphasizing realism implementation according to an embodiment of the present invention. In the broadcast processing, the defined target interest area can be uniformly enlarged, resulting in a larger display area for the target interest area within the entire display window, thereby ensuring that the viewer's gaze is fully focused on the content within the target display area.
[0107] Sub-step 40532: During the playback of the video stream, the edges of the target region of interest are highlighted.
[0108] Reference Figure 9 The diagram illustrates another implementation effect of emphasizing realism according to an embodiment of the present invention. Furthermore, since the edges of the target region of interest are clear geometric line segments, highlighting them can also achieve the purpose of emphasizing realism.
[0109] Sub-step 40533: Generate multiple display windows that float above the video stream image content based on the multiple regions of interest; the window size and position of the display windows correspond one-to-one with the multiple regions of interest whose director weight values meet preset conditions; and perform director processing on the image content of the multiple regions of interest through the multiple display windows respectively.
[0110] Reference Figure 10The diagram illustrates another implementation effect of content emphasis on realism provided by an embodiment of the present invention. A floating display window can be generated at the location of multiple regions of interest to emphasize the content within those regions. This method combines directing processing for regions of interest with the complete display of the video stream.
[0111] Sub-step 40534: Divide the current playback window into a first display area and a second display area that are independent of each other; play the student-end video stream through the first display area, and perform directing processing on the content of the interest area in the student-end video stream whose directing weight value meets the preset conditions; play the teacher-end video stream through the second display area, and perform directing processing on the content of the interest area in the teacher-end video stream whose directing weight value meets the preset conditions.
[0112] Furthermore, in this embodiment of the invention, considering that the acquired video streams come from multiple terminals, the video streams from the student's end and the teacher's end can be displayed in a split-screen format. Specifically, the playback window is divided into a first display area and a second display area that are independent of each other.
[0113] During the playback of the student-end video stream through the first display area, the content of the interest area in the student-end video stream that meets the preset director weight value is subject to director processing; at the same time, the teacher-end video stream is played through the second display area, and the content of the interest area in the teacher-end video stream that meets the preset director weight value is subject to director processing.
[0114] Sub-step 40535: During the process of displaying the student-end video stream through the first display device, the content of the interest region in the student-end video stream whose director weight value meets the preset conditions is subjected to director processing; during the process of displaying the teacher-end video stream through the second display device, the content of the interest region in the teacher-end video stream whose director weight value meets the preset conditions is subjected to director processing.
[0115] Similarly, in comparison sub-step 40534, video streams can be played separately through playback windows corresponding to the number of video streams, and content in interest areas whose director weight values meet preset conditions can be displayed. Specifically, during the process of displaying student-end video streams through the first display device, content in interest areas of student-end video streams whose director weight values meet preset conditions is subject to director processing; during the process of displaying teacher-end video streams through the second display device, content in interest areas of teacher-end video streams whose director weight values meet preset conditions is subject to director processing.
[0116] In summary, the image processing method provided by this invention first acquires the original video stream; then, it extracts regions of interest (ROIs) from the image frames of the video stream according to preset extraction rules, obtaining multiple ROIs; next, it inputs the content of the multiple ROIs extracted from the image frames into a content processing model to obtain a director weight value for each ROI; and during the playback of the video stream, it performs director processing on the image content within the target ROI region whose director weight value meets preset conditions. For the director processing of the entire video stream, instead of calculating all the content of the entire video frame, it prioritizes dividing multiple different ROIs according to the director weight values, and only performs analysis and calculation on key content within the ROI regions, eliminating the calculation process for non-ROI regions, reducing the computational load of image data, and making the application effect of director processing better.
[0117] Reference Figure 11 This diagram illustrates the functional components of an image processing apparatus according to an embodiment of the present invention. The apparatus includes:
[0118] The video stream acquisition module 501 is used to acquire the raw video stream.
[0119] The region of interest acquisition module 502 is used to extract regions of interest from the image frames of the video stream according to preset extraction rules, thereby obtaining multiple regions of interest.
[0120] The director weight value determination module 503 is used to input the content of multiple interest regions extracted from the image frame into the content processing model to obtain a director weight value for each interest region.
[0121] The directing processing execution module 504 is used to perform directing processing on the image content within the target interest area whose directing weight value meets preset conditions during the playback of the video stream.
[0122] Optionally, the device further includes:
[0123] The size adjustment module is used to adjust the size of the multiple regions of interest extracted from the image frame before inputting the content into the content processing model, so as to adjust the screen size of each of the multiple regions of interest to the same preset screen size.
[0124] Optionally, the region of interest acquisition module 502 further includes:
[0125] An object recognition submodule is used to perform object recognition on the image frame and obtain an object recognition box.
[0126] The region of interest generation submodule is used to select a portion of the area containing the object recognition box as the region of interest in response to the selection operation of the object recognition box.
[0127] Optionally, the region of interest acquisition module 502 further includes:
[0128] The first frame region of interest extraction submodule is used to extract the region of interest from the first image frame in the video stream according to a preset extraction rule, so as to obtain the region of interest of the first image frame.
[0129] The complete region of interest extraction submodule is used to map the contour position of the region of interest in the image frame at the first moment to the image frames at all other moments in the video stream, so as to obtain the region of interest in each image frame.
[0130] The video stream region of interest extraction submodule is used to extract regions from the student-side video stream to obtain multiple regions of interest including image frames from the student-side video stream, and to extract regions from the teacher-side video stream to obtain multiple regions of interest including image frames from the teacher-side video stream.
[0131] Optionally, the directing processing execution module 504 further includes:
[0132] The score sequence generation submodule is used to sort the regions of interest according to the director's weight value and obtain the score sequence of the regions of interest.
[0133] The target interest region determination submodule is used to determine the interest region whose director weight value is greater than the preset weight threshold as the target interest region when it is determined in the interest region score sequence that the director weight value of the interest region is greater than the preset weight threshold.
[0134] The emphasis is placed on the execution submodule, which is used to perform broadcasting processing on the image content within the target area of interest.
[0135] Optionally, the target region of interest determination submodule may further include:
[0136] The content event feature determination unit is used to obtain the content event type of the image frame within the region of interest, and the event priority corresponding to the content event type; the content event type is used to characterize the features of events occurring within the target region of interest.
[0137] The target region of interest determination unit is used to determine the region of interest with the highest event type priority among all the regions of interest whose director weight value is greater than the weight threshold, and to perform director processing on the image content within the target region of interest.
[0138] Optionally, the emphasis display execution submodule may further include:
[0139] A magnification center determination unit is used to take the geometric center of the target region of interest as the target magnification center;
[0140] The image magnification execution unit is used to, during the playback of the video stream, uniformly magnify the image content in the target region of interest to a preset size within a first duration, with the target magnification center as the reference, for display.
[0141] Optionally, the emphasis display execution submodule may further include:
[0142] The highlighting execution unit is used to highlight the edges of the target region of interest during the playback of the video stream.
[0143] Optionally, the emphasis display execution submodule may further include:
[0144] The floating window generation unit is used to generate multiple display windows that float above the video stream image content based on multiple regions of interest; the window size and position of the display windows correspond one-to-one with the multiple regions of interest whose director weight values meet preset conditions.
[0145] The multi-window display execution unit is used to perform broadcasting processing on the image content of multiple regions of interest through multiple display windows respectively.
[0146] Optionally, the emphasis display execution submodule may further include:
[0147] The display window partitioning unit is used to divide the current playback window into a first display area and a second display area that are independent of each other.
[0148] The first emphasis display execution unit is used to play the student terminal video stream through the first display area and perform directing processing on the content of the interest area in the student terminal video stream whose directing weight value meets the preset conditions;
[0149] The second emphasis display execution unit is used to play the teacher-side video stream through the second display area, and to perform directing processing on the content of the interest area in the teacher-side video stream whose directing weight value meets the preset conditions.
[0150] Optionally, the emphasis display execution submodule may further include:
[0151] The third emphasis is on the display execution unit, which is used to perform directing processing on the content of the interest area in the student terminal video stream whose directing weight value meets the preset conditions during the process of displaying the student terminal video stream through the first display device;
[0152] The fourth emphasis is on the display execution unit, which is used to perform directing processing on the content of the interest area in the teacher's video stream whose directing weight value meets the preset conditions during the process of displaying the teacher's video stream through the second display device.
[0153] Optionally, the video stream acquisition module 501 further includes:
[0154] The video stream acquisition submodule is used to acquire student-side video streams from the student's location and teacher-side video streams from the teacher's location.
[0155] In summary, the image processing apparatus provided by this invention first acquires the original video stream; then, it extracts regions of interest (ROIs) from the image frames of the video stream according to preset extraction rules, obtaining multiple ROIs; next, it inputs the content of the multiple ROIs extracted from the image frames into a content processing model to obtain a director weight value for each ROI; and during the playback of the video stream, it performs director processing on the image content within the target ROI region whose director weight value meets preset conditions. For the director processing of the entire video stream, instead of calculating all the content of the entire video frame, it prioritizes dividing multiple different ROIs according to the director weight values, and only performs analysis and calculation on key content within the ROI regions, eliminating the calculation process for non-ROI regions, reducing the computational load of image data, and making the application effect of director processing better.
[0156] Figure 12 This is a block diagram illustrating an electronic device 600 according to an exemplary embodiment. For example, the electronic device 600 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0157] Reference Figure 12 The electronic device 600 may include one or more of the following components: a processing component 602, a memory 604, a power supply component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.
[0158] Processing component 602 typically controls the overall operation of electronic device 600, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 602 may include one or more modules to facilitate interaction between processing component 602 and other components. For example, processing component 602 may include a multimedia module to facilitate interaction between multimedia component 608 and processing component 602.
[0159] Memory 604 is used to store various types of data to support the operation of electronic device 600. Examples of such data include instructions for any application or method operating on electronic device 600, contact data, phonebook data, messages, pictures, multimedia, etc. Memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0160] Power supply component 606 provides power to various components of electronic device 600. Power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 600.
[0161] Multimedia component 608 includes a screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 608 includes a front-facing camera and / or a rear-facing camera. When the electronic device 600 is in an operating mode, such as a shooting mode or a multimedia mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0162] Audio component 610 is used to output and / or input audio signals. For example, audio component 610 includes a microphone (MIC) used to receive external audio signals when electronic device 600 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 604 or transmitted via communication component 616. In some embodiments, audio component 610 also includes a speaker for outputting audio signals.
[0163] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0164] Sensor assembly 614 includes one or more sensors for providing state assessments of various aspects of electronic device 600. For example, sensor assembly 614 can detect the on / off state of electronic device 600, the relative positioning of components such as the display and keypad of electronic device 600, changes in position of electronic device 600 or a component of electronic device 600, the presence or absence of user contact with electronic device 600, orientation or acceleration / deceleration of electronic device 600, and temperature changes of electronic device 600. Sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 614 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0165] Communication component 616 facilitates wired or wireless communication between electronic device 600 and other devices. Electronic device 600 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 616 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 616 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0166] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement an image processing method provided in an embodiment of the present invention.
[0167] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by a processor 620 of an electronic device 600 to perform the above-described method. For example, the non-transitory storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0168] Figure 13This is a block diagram illustrating an electronic device 700 according to an exemplary embodiment. For example, the electronic device 700 may be provided as a server. (Refer to...) Figure 13 The electronic device 700 includes a processing component 722, which further includes one or more processors, and memory resources represented by a memory 732 for storing instructions, such as application programs, that can be executed by the processing component 722. The application programs stored in the memory 732 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 722 is configured to execute instructions to perform an image processing method provided in an embodiment of the present invention.
[0169] Electronic device 700 may also include a power supply component 726 configured to perform power management of electronic device 700, a wired or wireless network interface 750 configured to connect electronic device 700 to a network, and an input / output (I / O) interface 758. Electronic device 700 may operate on an operating system stored in memory 732, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.
[0170] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.
[0171] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, The method, applied to an AI-based smart classroom broadcasting system, includes: Obtain the raw video stream; According to preset extraction rules, the image frames of the video stream are subjected to region of interest extraction to obtain multiple regions of interest; The size of each of the multiple regions of interest is adjusted to bring the screen size of each region of interest to the same preset screen size, thereby obtaining multiple regions of interest with the same size. The content of the multiple interest regions of the same size is input into the content processing model for evaluation to obtain the broadcast weight value for each interest region. The multiple batch dimensions of the content processing model correspond to the content of the multiple interest regions of the same size. During the playback of the video stream, the image content within the target area of interest whose director weight value meets the preset conditions is subjected to director processing. The director processing is used to play the image content and / or panoramic view of the target area of interest on the display device.
2. The method according to claim 1, characterized in that, The step of extracting regions of interest (ROIs) from image frames of the video stream according to preset extraction rules to obtain multiple ROIs includes: Perform object recognition on the image frame to obtain object recognition boxes; In response to the selection operation of the object recognition box, the area where part of the object recognition box is located is selected as the region of interest.
3. The method according to claim 2, characterized in that, The number of regions of interest is directly proportional to the density of the object recognition box.
4. The method according to claim 1, characterized in that, The step of extracting regions of interest (ROIs) from image frames of the video stream according to preset extraction rules to obtain multiple ROIs includes: According to the preset extraction rules, the region of interest is extracted from the first image frame in the video stream to obtain the region of interest of the first image frame. The contour position of the region of interest in the image frame at the first moment is mapped to the image frames at all other moments in the video stream to obtain the region of interest for each image frame.
5. The method according to claim 1, characterized in that, The broadcast weight value of the region of interest is directly proportional to the broadcast priority of the image content within the region of interest.
6. The method according to claim 1, characterized in that, The step of performing broadcast processing on the image content within the target interest region whose broadcast weight value meets preset conditions includes: Based on the director's weight value, the regions of interest are sorted to obtain a score sequence for each region of interest. In the interest region score sequence, if it is determined that the director weight value of an interest region is greater than a preset weight threshold, the interest region whose director weight value is greater than the weight threshold is determined as the target interest region. The image content within the target area of interest is then processed for broadcasting.
7. The method according to claim 6, characterized in that, In the interest region score sequence, if it is determined that the director weight values of multiple interest regions are all greater than the weight threshold, the method further includes: Obtain the content event type of the image frame within the region of interest, and the event priority corresponding to the content event type; the content event type is used to characterize the features of events occurring within the target region of interest; Among all the interest regions whose director weight value is greater than the weight threshold, the interest region with the highest priority of the event type is determined as the target interest region, and the image content within the target interest region is processed by the director.
8. The method according to claim 1, characterized in that, The step of performing broadcasting processing on the image content within the target area of interest includes: Use the geometric center of the target region of interest as the target magnification center; During the playback of the video stream, the image content in the target region of interest is uniformly enlarged to a preset size within a first duration, based on the target magnification center.
9. The method according to claim 1, characterized in that, The step of performing broadcasting processing on the image content of the target region of interest also includes: During the playback of the video stream, the edges of the target region of interest are highlighted.
10. The method according to claim 1, characterized in that, If it is determined that the director weight values of multiple regions of interest meet preset conditions, the image content within the regions of interest whose director weight values meet the preset conditions is subjected to director processing, including: Based on the multiple regions of interest, multiple display windows are generated that float above the video stream image content; the window size and position of the display windows correspond one-to-one with the multiple regions of interest whose director weight values meet preset conditions. The image content of multiple interest areas is directed and processed through multiple display windows respectively.
11. The method according to claim 1, characterized in that, The acquisition of the original video stream includes: Acquire student-side video streams from the perspective of the student's location, and teacher-side video streams from the perspective of the teacher's location; The step of extracting regions from image frames of the video stream according to preset extraction rules to obtain multiple regions of interest includes: Region extraction is performed on the student-side video stream to obtain multiple regions of interest (ROIs) including image frames from the student-side video stream, and region extraction is also performed on the teacher-side video stream to obtain multiple regions of interest including image frames from the teacher-side video stream.
12. The method according to claim 11, characterized in that, During the playback of the video stream, the image frames within the target interest region are subjected to directing processing based on the weight value, including: Divide the current playback window into two independent display areas: a first display area and a second display area. The student-side video stream is played through the first display area, and the content of the interest area in the student-side video stream whose director weight value meets the preset conditions is subjected to director processing. The teacher's video stream is played through the second display area, and the content of the interest area in the teacher's video stream whose director weight value meets the preset conditions is subject to director processing.
13. The method according to claim 11, characterized in that, During the playback of the video stream, the image frames within the target interest region are subjected to directing processing based on the weight value, including: During the process of displaying the student-end video stream through the first display device, the content of the interest region in the student-end video stream whose director weight value meets the preset conditions is subject to director processing; During the process of displaying the teacher's video stream through the second display device, the content of the interest area in the teacher's video stream whose director weight value meets the preset conditions is subject to director processing.
14. An image processing apparatus, characterized in that, The device includes: The video stream acquisition module is used to acquire the raw video stream; The region of interest (ROI) acquisition module is used to extract ROIs from the image frames of the video stream according to preset extraction rules to obtain multiple ROIs; and to adjust the size of each of the multiple ROIs to adjust the screen size of each of the multiple ROIs to the same preset screen size, thereby obtaining multiple ROIs of the same size. The director weight value determination module is used to input the content of the multiple interest regions of the same size into the content processing model for evaluation, and obtain the director weight value for each interest region, wherein the multiple batch dimensions of the content processing model correspond to the content of the multiple interest regions of the same size; The directing processing execution module is used to perform directing processing on the image content of the target interest area whose directing weight value meets preset conditions during the playback of the video stream. The directing processing is used to play the image content and / or panoramic view of the target interest area on the display device.
15. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the method as described in any one of claims 1 to 13.
16. A readable storage medium, characterized in that, When the instructions in the readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1 to 13.
Citation Information
Patent Citations
Panoramic video live broadcast method and system and computer readable storage medium
CN113099245A