Multi-person close-up cutting and splicing display method based on human body detection and storage medium
By performing human detection and coordinate correction on the image data obtained by the camera, the cropping and splicing display of multiple people close-ups in the conference scene is realized, and the problem of multi-person close-ups display and splicing encoding in the prior art is solved.
Patent Information
- Application Number
- CN202411911491.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to display close-ups of multiple participants in a conference scenario at the same time, and there is a lack of a method of encoding the entire close-up splicing screen.
The camera module obtains image data, recognizes the human body detection area and obtains coordinate information, corrects the coordinate information, so that the length and width ratio of the detection area is consistent with the display area, and displays the corrected detection area on the display area, realizing the cropping and splicing of close-ups of multiple people.
It realizes the close-up display of multiple participants in the conference scene at the same time, and the entire close-up splicing screen is encoded, solving the problem of close-up display and splicing encoding of multiple people in the prior art.
Smart Images

Figure CN120070827A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology, and particularly relates to a method and storage medium for cropping, splicing and displaying multiple-person close-ups based on human body detection. Background Art
[0002] Nowadays, human body recognition technology is very mature, and there are many related applications, which are also widely used in conference scenarios. In actual conference scenarios, the background of the conference room is not of great concern, and the main focus is on the status of the participants. By inputting the camera to capture the conference scene image, performing human body recognition on it, and returning the detected area coordinates, which are the positions of the participants in the image captured by the camera. Generally, only the people are simply framed by these coordinates, and in this way, the background of the conference will still be displayed. Further, the close-ups of the participants can be displayed through the detected area coordinates. However, in the case of multiple people, this requires separately cropping the human body area parts in the image and then splicing them together. Since the distances of the human bodies from the camera are different, the sizes of the detected areas are significantly different, and the splicing also requires strategy adjustment. And in the conference room scenario, there is often a need to send the image of the local camera to the remote end, which also has encoding requirements for the spliced image and cannot be only locally displayed. Summary of the Invention
[0003] In view of the above problems, the present application provides a method and storage medium for cropping, splicing and displaying multiple-person close-ups based on human body detection, which solves the problem that for the areas detected by existing human body detection and recognition, most of them only draw wire frames or display single-person close-ups in rotation, and there is no method that can meet the requirement of simultaneously displaying the close-ups of multiple participants in the conference scenario and encoding the entire close-up splicing image.
[0004] To achieve the above object, the inventor provides a method for cropping, splicing and displaying multiple-person close-ups based on human body detection, including:
[0005] Obtaining image data through a camera module;
[0006] Identifying the detection area where the human body is located in the image data and returning the coordinate information of the detection area in the image data;
[0007] Correcting the returned coordinate information so that the aspect ratio of the detection area is corrected to be consistent with the ratio of the display area;
[0008] Displaying the corrected detection area on the corresponding display area.
[0009] In some embodiments, the detection area is a rectangular area, and the coordinate information is the pixel positions of the upper left point and the lower right point of the detection area in the image data.
[0010] In some embodiments, the correction of the returned coordinate information to make the aspect ratio of the detection area consistent with that of the display area specifically includes the following steps:
[0011] Calculate the center point of the detection area;
[0012] Determine whether the size of the detection area is smaller than the minimum area;
[0013] If so, expand the detection area to the size of the minimum area according to the detected center point;
[0014] If not, expand the detection area according to a preset ratio.
[0015] In some embodiments, it further includes the following steps:
[0016] When expanding the detection area, if the expanded area is at the boundary, expand in the other direction.
[0017] In some embodiments, it further includes the following steps:
[0018] When the detection areas where the recognized human bodies are located exceed the preset number, merge the detection areas within the preset range into one area.
[0019] In some embodiments, the merging of the detection areas within the preset range into one area specifically includes the following steps:
[0020] Evenly divide the display window into a preset number of display areas;
[0021] Merge the detection areas whose center points are within the same display area into one detection area.
[0022] In some embodiments, it specifically includes the following steps:
[0023] Create an off-screen buffer and a window buffer through openGL;
[0024] Correct the coordinate information of the detection area in the off-screen buffer;
[0025] Copy the detection area with the corrected coordinate information to the display area in the window buffer for display.
[0026] In some embodiments, the copying of the detection area with the corrected coordinate information to the display area in the window buffer for display includes the following steps:
[0027] Set the corresponding number of display areas in the window buffer according to the number of recognized detection areas;
[0028] After correcting the detection area in the off-screen buffer, copy it to the corresponding display area in the window buffer;
[0029] Stitch all the display areas in the window cache into a window and then display it on the screen.
[0030] In some embodiments, the image data is video data in NV12 format.
[0031] Another technical solution is also provided, a storage medium storing a computer program, and when the computer program is run by a processor, it executes the steps of the method for cropping, splicing and displaying multiple people's close-ups based on human body detection as described above.
[0032] Different from the prior art, in the above technical solution, image data is obtained through a camera module, the detection area where the human body is located in the image data is identified, and the coordinate information of the detection area in the image data is obtained. After correcting the obtained coordinate information, the detection area with the corrected coordinate information is displayed on the corresponding display area. When identifying the human body in the image data, when multiple detection areas where multiple human bodies are located are identified, the coordinate information of the detection areas in the corresponding multiple data graphics is returned, and then the coordinate information is corrected so that the length-width ratio of each detection area is the same as that of the display area, so that when the detection area is displayed in the corresponding display area in the window, the sizes of the individual detection areas are the same; at the same time, the individual detection areas are displayed in the display area of the window, realizing the splicing display of the individual detection areas, and enabling the close-up display of multiple human bodies.
[0033] The above related description of the invention content is only an overview of the technical solution of this application. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of this application, and then implement it according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features and advantages of this application more easily understood, the following is described in conjunction with the specific embodiments of this application and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings are only used to illustrate the principles, implementation methods, applications, features and effects of the specific embodiments of the present invention and other related contents, and should not be considered as a limitation to this application.
[0035] In the accompanying drawings of the specification:
[0036] Figure 1 It is a flow chart of a method for cropping, splicing and displaying multiple people's close-ups based on human body detection as described in the specific embodiment;
[0037] Figure 2 It is a flow chart of the coordinate correction as described in the specific embodiment;
[0038] Figure 3A schematic diagram of a process for copying a detection area from an off-screen buffer to a window buffer as described in the specific implementation manner;
[0039] Figure 4 A schematic diagram of an application module of the multi-person close-up cropping, stitching, and display method based on human detection as described in the specific implementation manner;
[0040] Figure 5 Another schematic diagram of a process for the multi-person close-up cropping, stitching, and display method based on human detection as described in the specific implementation manner;
[0041] Figure 6 A schematic diagram of a structure of the storage medium as described in the specific implementation manner.
[0042] The descriptions of the reference numerals involved in the above-mentioned drawings are as follows:
[0043] 610. Storage medium,
[0044] 620. Processor. Specific implementation manner
[0045] To illustrate in detail the possible application scenarios, technical principles, specific implementable solutions, achievable purposes and effects, etc. of the present application, the following will be described in detail in conjunction with the listed specific examples and with reference to the accompanying drawings. The examples described herein are only used to more clearly illustrate the technical solutions of the present application, so they are only examples and cannot be used to limit the protection scope of the present application.
[0046] Referring to "embodiment" in this article means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The term "embodiment" that appears in various positions in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in the present application, as long as there is no technical contradiction or conflict, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.
[0047] Unless otherwise defined, the meanings of the technical terms used in this article are the same as those generally understood by those skilled in the technical field to which the present application belongs; the use of relevant terms in this article is only for describing specific embodiments and is not intended to limit the present application.
[0048] In the description of the present application, the term "and / or" is an expression used to describe the logical relationship between objects, indicating that there can be three relationships. For example, A and / or B means: there is A, there is B, and there is both A and B at the same time. In addition, the character " / " in this article generally represents an "or" logical relationship between the associated objects before and after.
[0049] In this application, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantitative, primary-secondary, or sequential relationship between these entities or operations.
[0050] Without further limitation, in this application, the open-ended expressions such as "comprising", "including", "having", or other similar expressions used in a statement are intended to cover non-exclusive inclusion. These expressions do not exclude the possibility that there may be additional elements in the process, method, or product that includes the said elements. Thus, in a process, method, or product that includes a series of elements, it may include not only those defined elements, but also other elements not explicitly listed, or elements inherent to such a process, method, or product.
[0051] Similar to the understanding in the Examination Guidelines, in this application, expressions such as "greater than", "less than", "exceeding", etc. are understood not to include the recited number; expressions such as "above", "below", "within", etc. are understood to include the recited number. In addition, in the description of the embodiments of this application, the meaning of "a plurality of" is two or more (including two). Similar expressions related to "multiple", such as "multiple groups", "multiple times", etc., are understood in the same way, unless otherwise specifically defined.
[0052] In the description of the embodiments of this application, the spatially related expressions used, such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "perpendicular", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the specific embodiment or the drawings. It is only for the convenience of describing the specific embodiments of this application or for the reader's understanding, and does not indicate or imply that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, it should not be construed as a limitation on the embodiments of this application.
[0053] The processor described in the embodiments of the present application can be implemented by hardware, firmware, software, or a combination thereof, and can use circuits, one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, microprocessors, or at least one of the above, and also includes other physical, biological, or chemical structures that can achieve functions similar to or equivalent to those of the above-listed processors, such as biological neurons, quantum computing units, DNA computing units, etc., so that the processor can execute some steps, all steps, or any combination of the steps mentioned in the computer programs or methods involved in the various embodiments of the present application.
[0054] The computer programs involved in the embodiments can be stored in a computer-readable storage medium, which includes but is not limited to magnetic disks, magnetic tapes, magnetic cards, floppy disks, flash memories, optical discs, optical cards, read-only memories (ROMs), random access memories (RAMs), erasable programmable ROMs (EPROMs), and electrically erasable programmable ROMs (EEPROMs), etc. It also includes other biological, physical, or chemical structures that can achieve the same or equivalent functions as the above-listed storage media, such as units with information storage capabilities like DNA, RNA, proteins, etc. In a specific embodiment, the storage medium involved can be one of the above medium types or a combination of the above medium types. In different embodiments, the computer programs involved in the embodiments can be stored centrally in a single medium or distributedly in multiple media. The memory containing the computer-readable storage medium can be a non-volatile memory or a random access memory. These computer-readable storage media can be built into the device or connected to the device involved in the embodiment as an external device or a part of an external device. In some embodiments, the memory with the computer-readable storage medium is deployed locally; in other embodiments, a scheme of deploying the memory away from the processor can also be adopted, such as a network-attached memory accessed via an RF circuit or an external port and a communication network, where the communication network can be the Internet, one or more intranets, local area networks (LANs), wide area wireless networks (WLANs), storage area networks (SANs), etc., or a suitable combination thereof, as long as the computer device can access the memory. In addition, the computer programs involved in the embodiments can be stored in plaintext / ciphertext form or designed as training data and integrated and reorganized implicitly and saved in the parameter states of a deep neural network or other machine learning models through model training.
[0055] Please refer to Figure 1 , this embodiment provides a method for cropping, splicing, and displaying multiple-person close-ups based on human body detection, including:
[0056] Step S110: Obtain image data through a camera module;
[0057] Step S120: Identify the detection area where the human body is located in the image data and return the coordinate information of the detection area in the image data;
[0058] Step S130: Correct the returned coordinate information so that the aspect ratio of the detection area is corrected to be the same as that of the display area;
[0059] Step S140: Display the corrected detection area on the corresponding display area.
[0060] The image data is acquired through a camera module, the detection area where the human body is located in the image data is recognized, and the coordinate information of the detection area in the image data is obtained. After the obtained coordinate information is corrected, the detection area with the corrected coordinate information is displayed on the corresponding display area. When recognizing the human body in the image data, when multiple detection areas where human bodies are located are recognized, the coordinate information of the detection areas in the corresponding multiple data graphics is returned, and then the coordinate information is corrected so that the length-width ratio of each detection area is consistent with the length-width ratio of the display area, so that when the detection area is displayed in the corresponding display area in the window, the sizes of the respective detection areas are the same; at the same time, the respective detection areas are displayed in the display area of the window, realizing the tiled display of the respective detection areas, and enabling the close-up display of multiple human bodies.
[0061] In some embodiments, the detection area is a rectangular area, and the coordinate information is the pixel positions of the upper left point and the lower right point of the detection area in the image data.
[0062] Generally, the display area is rectangular, so the detection area is also set as a rectangular area; in other embodiments, the detection area can also be set as a circle or an ellipse, etc. When the detection area is a rectangular area, the coordinate information of the returned detection area is the pixel positions of the upper left point and the lower right point of the rectangular area in the entire image data, and a rectangular area can be determined through these two points. In other embodiments, the coordinate information can also include the two pixel positions of the upper right point and the lower left point, or the four pixel positions of the upper left point, the upper right point, the lower left point, and the lower right point.
[0063] In some embodiments, the correction of the returned coordinate information to make the length-width ratio of the detection area consistent with the ratio of the display area specifically includes the following steps:
[0064] Calculate the center point of the detection area;
[0065] Determine whether the size of the detection area is smaller than the minimum area;
[0066] If so, expand the detection area to the size of the minimum area according to the detected center point;
[0067] If not, expand the detection area according to a preset ratio.
[0068] Correction to the coordinates returned by the human body detection module. Since the distance between the human body and the camera is different, the length and width of the detected area are often different. In order to avoid excessive compression or stretching of the final display, the length and width ratio of the detection coordinates need to be corrected to make it consistent with the display window ratio. At the same time, in order to avoid the person being too far away from the camera, resulting in the detection area being displayed on the window too enlarged, the picture is blurred, and the mosaic is obvious, a minimum rectangular area size is set. If the detection area is smaller than this area, it will be expanded. Figure 2 As shown in the figure, the process of coordinate correction is as follows:
[0069] 1. After obtaining human body detection, first calculate the center point of the detection area based on its coordinate information;
[0070] 2. Set the minimum area size to minWidth*minHeight. If the detection area is smaller than the minimum area, it will be expanded to minWidth*minHeight based on the center point of the detection area; if the detection area is larger than the minimum area, the height remains unchanged and the width is expanded to the preset ratio, such as width: height = 16:9. The width can also be used as a reference when correcting the coordinates, keeping the width unchanged, and select according to actual needs.
[0071] In some embodiments, the following steps are also included:
[0072] When the detection area is expanded, if the expanded area is at the boundary, it is expanded in another direction.
[0073] If the expanded area is at the boundary, expand it in the other direction and try to keep the proportion unchanged.
[0074] In some embodiments, the following steps are also included:
[0075] When the number of detection areas where the identified human body is located exceeds a preset number, the detection areas within the preset range are merged into one area.
[0076] For multi-person scenes, when the screen space is limited and there are too many people, it is inevitable that some people are close to each other. At this time, there are actually a lot of repeated scenes between the two. If they are displayed separately, the effect will not be good. When more than a certain number of people are detected in the picture, it is necessary to merge the detection areas that are close to each other, and recalculate an area coordinate to include all overlapping or close detection areas.
[0077] In some embodiments, merging the detection areas within the preset range into one area specifically includes the following steps:
[0078] Divide the display window evenly into a preset number of display areas;
[0079] Merge the detection regions with their centers within the same display area into one detection region.
[0080] It is possible to count the number of display areas with the largest support in the display window as N, then evenly divide the entire display window into N display areas, put the detection regions with their centers in the same display area into a table, and perform merging. This can save the system overhead of traversing and sorting the detection regions to calculate which ones are closer. When N is set relatively large, the evenly divided display areas are relatively small, and the detection regions need to be relatively close to meet the requirements, so that two detection regions that are far apart will not be merged either.
[0081] In some embodiments, it specifically includes the following steps:
[0082] Create an off-screen buffer and a window buffer through openGL;
[0083] Correct the coordinate information of the detection region in the off-screen buffer;
[0084] Copy the detection region with corrected coordinate information to the display area in the window buffer for display.
[0085] The display rendering is mainly implemented through openGL, and a virtual off-screen buffer and a window buffer of the actual screen window will be applied. As Figure 3 shown, the original data read from the camera module is first saved in the off-screen buffer, and then according to the previously adjusted region coordinates, the corresponding detection region is copied from the off-screen buffer to the specified region in the window buffer, which realizes the cropping and splicing of the human detection region. Finally, the front and back buffer areas of the window are swapped, and the final result will be rendered on the screen.
[0086] In some embodiments, the step of copying the detection region with corrected coordinate information to the display area in the window buffer for display includes the following steps:
[0087] Set the corresponding number of display areas in the window buffer according to the number of detected detection regions;
[0088] After correcting the detection region in the off-screen buffer, copy it to the corresponding display area in the window buffer;
[0089] Stitch all the display areas in the window buffer into a window and display it on the screen.
[0090] The detected area of the human body in the off-screen cache is expanded and then copied to the specified display area of the window. During the copying process, OpenGL will automatically perform scaling, and the copying process supports cropping of the original picture. Multiple display areas can be set in the window cache area. When multiple different detection areas are detected, the entire window cache area is divided into multiple display areas according to requirements, and different detection areas are copied to the display areas. All the display areas form an entire window, and the splicing effect is achieved when displayed. The display areas here can be customized according to requirements, which is relatively flexible and can achieve different display layouts.
[0091] In some embodiments, the image data is video data in NV12 format. The NV12 format is a video coding format and a variant of the YUV color space. The NV12 format is defined by Intel and is natively supported mainly on Intel hardware platforms. Its characteristic is that in the memory arrangement, the Y component is arranged continuously first, and then the U and V components are arranged alternately. This arrangement makes the NV12 format have high efficiency when processing video data.
[0092] In some embodiments, a method for cropping, splicing and displaying multiple people's close-ups based on human body detection is provided. As Figure 4 shown, the entire application module can be divided into four parts: a camera module, a human body detection module, a coordinate correction module, and a display rendering module.
[0093] Camera module: mainly provides the camera picture, and the camera needs to call back data in NV12 format. After obtaining the data, it is sent to the human body detection module for detecting the coordinates of the human body detection area.
[0094] Human body detection module: identifies the position of the detection area in the picture and returns the coordinates of the detection area in the image. The coordinates are the pixel positions of the upper left point and the lower right point of a rectangular area in the entire image. Through these two points, a rectangular area can be determined, and the target of human body detection is in the area. When multiple people are detected, the positions of multiple rectangular coordinates are returned.
[0095] Coordinate correction module: corrects the coordinates returned by the human body detection module. Since the distances of the human body from the camera are different, the detected area lengths and widths are often different. To avoid excessive compression or stretching of the final display, it is necessary to correct the length-width ratio of the detection coordinates to make it consistent with the display window ratio. At the same time, to avoid the situation where the person is too far from the camera, resulting in the detection area being too enlarged on the window, the picture being blurred, and the mosaic being obvious, a minimum rectangular area size is set, and if the detection area is smaller than this area, it is expanded. The specific correction process is as shown in the figure;
[0096] 1. After obtaining the human body detection area, first calculate the center point of the detection area;
[0097] 2. Set the minimum area size to minWidth*minHeight. If the detection area is smaller than the minimum area, it will be expanded to minWidth*minHeight based on the center point of the detection area; if the detection area is larger than the minimum area, the height remains unchanged and the width is expanded until width:height = 16:9. The width can also be used as a reference when correcting the coordinates, keeping the width unchanged, and it can be selected according to actual needs.
[0098] 3. If the expanded area is at the boundary, expand in the other direction and try to keep the proportion unchanged.
[0099] The above is the correction and expansion process of a single detection window. For multi-person screens, when the screen space is limited and there are too many people, it is inevitable that some people are close to each other. At this time, there are actually a lot of repeated screens between the two. If they are displayed separately, the effect will not be good. When more than a certain number of people are detected in the screen, it is necessary to merge the detection areas that are close to each other, and recalculate the coordinates of an area to include all overlapping or close detection areas. The maximum number of windows supported for display can be counted as N, and then the entire display area can be divided into N blocks on average. The detection areas with the center points in the same display block are put into the table and merged. This can save the system overhead of traversing and sorting the detection areas to calculate which ones are closer. When N is set to a large value, the display area divided equally is smaller, and the detection areas must be closer to meet the requirements, so that two detection areas that are far apart will not be merged.
[0100] Display rendering module: Display rendering is mainly implemented through openGL, which will apply for a virtual off-screen buffer and an actual screen window buffer. The original data read from the camera is first saved in the off-screen buffer, and then the corresponding detection area is copied from the off-screen buffer to the specified area of the screen window buffer according to the previously adjusted area coordinates, which realizes the cropping and splicing of the human body detection area. Finally, the front and back buffer areas of the window are exchanged, and the final result will be rendered on the screen. The copying process is shown in the figure;
[0101] The area detected by the human body in the off-screen buffer is expanded and copied to the specified display area of the window. OpenGL will automatically scale during the copying process, and the copying process supports the cropping of the original image. The window buffer area can set multiple display areas. When multiple different detection areas are detected, the entire window buffer area is divided into multiple display areas according to requirements, and different detection areas are copied to the display areas. All display areas form a whole window, and the display achieves a splicing effect. The display area here can be customized according to requirements, which is more flexible and can achieve different display layouts.
[0102] Specifically, the process of the multi-person close-up cropping and stitching display method based on human body detection is as follows Figure 5 as shown, including the following steps:
[0103] Initialize the detection module;
[0104] Turn on the camera module and set the camera data callback;
[0105] After the camera obtains each frame of raw data, the human body detection module performs detection;
[0106] When a human body is detected, the detection information is called back;
[0107] Return the coordinates of the detection area;
[0108] Judge whether the number of detection areas exceeds the preset maximum window number;
[0109] If so, merge the detection coordinates in the same area, and then correct the coordinates of the detection area;
[0110] If not, correct the coordinates of the detection area;
[0111] After correction, the backrest detection area is sent to the display window for swap buffer display;
[0112] When finished, turn off the camera and destroy all modules accordingly.
[0113] The close-ups of multiple people can be displayed simultaneously on one window, and this window can also be encoded and sent to the remote end separately, and the layout of the picture can also be adjusted flexibly.
[0114] Please refer to Figure 6 , in another embodiment, a storage medium 610 stores a computer program, and when the computer program is run by a processor 620, it executes the steps of the multi-person close-up cropping and stitching display method based on human body detection as described above.
[0115] Image data is obtained through a camera module, the detection area where the human body is located in the image data is recognized, and the coordinate information of the detection area in the image data is obtained. After correcting the obtained coordinate information, the detection area with the corrected coordinate information is displayed on the corresponding display area. When recognizing the human body in the image data, when multiple detection areas where the human body is located are recognized, the coordinate information of the detection areas in the corresponding multiple data graphics is returned, and then the coordinate information is corrected so that the length and width ratio of each detection area is the same as that of the display area, so that when the detection area is displayed in the corresponding display area in the window, the sizes of the individual detection areas are the same; at the same time, the individual detection areas are displayed in the display area of the window, and the splicing display of the individual detection areas is realized, and the close-up display of multiple human bodies can be realized.
[0116] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of the present application, the patent protection scope of the present application cannot be limited thereby. Any technical solution obtained by equivalent structure or equivalent process substitution or modification based on the substantial concept of the present application and using the content recorded in the text and drawings of the specification of the present application, as well as directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, etc., are all included in the patent protection scope of the present application.
Claims
1. A method for cropping and splicing close-up images of multiple people based on human body detection, characterized in that: include: Acquire image data through the camera module; Identify the detection area where the human body is located in the image data, and return the coordinate information of the detection area in the image data; Correct the returned coordinate information so that the aspect ratio of the detection area is corrected to be consistent with the aspect ratio of the display area; The corrected detection area is displayed on the corresponding display area.
2. The method for cropping, splicing and displaying multiple close-up images based on human body detection according to claim 1, characterized in that: The detection area is a rectangular area, and the coordinate information is the upper left pixel position and the lower right pixel position of the detection area in the image data.
3. The method for cropping, splicing and displaying multiple close-up images based on human body detection according to claim 2, characterized in that: The step of correcting the returned coordinate information so that the aspect ratio of the detection area is corrected to be consistent with the aspect ratio of the display area specifically includes the following steps: Calculate the center point of the detection area; Determine whether the size of the detection area is smaller than the minimum area; If yes, then the detection area is expanded to the size of the minimum area based on the center point of the detection; If not, the detection area is expanded according to a preset ratio.
4. The method for cropping, splicing and displaying multiple close-up images based on human body detection according to claim 3 is characterized in that: The following steps are also included: When the detection area is expanded, if the expanded area is at the boundary, it is expanded in another direction.
5. The method for cropping, splicing and displaying multiple close-up images based on human body detection according to claim 1, characterized in that: The following steps are also included: When the number of detection areas where the identified human body is located exceeds a preset number, the detection areas within the preset range are merged into one area.
6. The method for cropping, splicing and displaying multiple close-up images based on human body detection according to claim 5, characterized in that: The step of merging the detection areas within the preset range into one area specifically comprises the following steps: Divide the display window evenly into a preset number of display areas; The detection areas whose center points are within the same display area are merged into one detection area.
7. The method for cropping, splicing and displaying close-up images of multiple people based on human body detection according to claim 1, characterized in that: The specific steps include: Create off-screen cache and window cache through openGL; Correct the coordinate information of the detection area in the off-screen cache; The detection area after the coordinate information is corrected is copied to the display area in the window buffer for display.
8. The method for cropping, splicing and displaying multiple close-up images based on human body detection according to claim 7, characterized in that: The step of copying the detection area after the coordinate information is corrected to the display area in the window cache for display comprises the following steps: According to the number of identified detection areas, a corresponding number of display areas are set in the window cache; After the detection area is corrected in the off-screen buffer, it is copied to the corresponding display area in the window buffer; All display areas in the window cache are stitched into one window and displayed on the screen.
9. The method for cropping, splicing and displaying close-up images of multiple people based on human body detection according to claim 1, characterized in that: The image data is video data in NV12 format.
10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the multi-person close-up cropping, splicing and display method based on human body detection as described in any one of claims 1 to 9 are executed.