Method and system for detecting reflective vest of personnel in large scene
Through the collaborative work of the main camera and the ball machine, people can be captured and tracked in real time, centered amplification processing and segmented model judgment, solving the problem of insufficient accuracy of reflective clothing detection in large scenarios, and achieving high accuracy and low missed false alarm detection effects.
Patent Information
- Application Number
- CN202411886126.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-20
AI Technical Summary
In large scenarios, traditional target detection methods are difficult to accurately determine whether the person is wearing reflective clothes, especially when the occlusion, overlap and side reflective clothes account for a small proportion, resulting in missed and false alarms.
Use the main camera to capture video streams in real time, identify and track people, and predict their target locations. Then, the centered magnification is performed by the associated ball machine, and the upper body and reflective clothing of the person are divided using a segmentation model to determine whether to wear reflective clothing.
It improves the accuracy of the wear detection of reflective clothing in small targets (personnel) in large scenarios, reduces the rate of missed and false alarms, and is suitable for reflective clothing detection in various large scenarios.
Smart Images

Figure CN120107876A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a method and system for detecting reflective clothing of people in a large scene. Background Art
[0002] In large scenes such as ports and airports, it is crucial to ensure the safe operation of personnel. Among them, ensuring that staff members wear reflective clothing correctly is one of the key links in safety management. However, existing personnel detection and recognition technologies have many limitations when applied to the detection of reflective clothing for personnel in large scenes. Such conventional methods have the following problems. First, since the proportion of personnel in the image is small, traditional target detection methods are difficult to accurately determine whether personnel are wearing reflective clothing. Secondly, reflective clothing is worn over regular clothing, and there are many types of regular clothing, which requires traditional target detection methods to continuously stack samples to cope with various types of clothing. This method is not only inefficient and difficult to maintain, but also difficult to meet accuracy requirements. Finally, in complex scenes, there are often occlusions and overlaps between personnel, and the proportion of reflective clothing on the side of personnel is small. In this case, traditional algorithms are difficult to accurately identify the reflective clothing of each individual, resulting in missed reports and false reports.
[0003] Therefore, how to improve the detection accuracy of people's reflective clothing is a technical problem that needs to be solved at present. Summary of the invention
[0004] In order to at least solve the technical problems existing in the above-mentioned background technology, the present invention provides a method, system, electronic device, computer storage medium and computer program product for detecting reflective clothing of people in a large scene.
[0005] A first aspect of the present invention provides a method for detecting reflective clothing of a person in a large scene, comprising the following steps: S1. Use a main camera to capture a video stream in real time, identify persons appearing in the video stream, and detect, track and predict each of the persons to obtain a predicted target position of each of the persons; S2, determining a number of associated ball cameras according to the target position, and controlling at least one of the associated ball cameras to perform centering and zooming processing on the person in the ball camera screen; S3, in the ball camera image after the center enlargement processing, using the segmentation model to segment the upper body and reflective clothing of the person respectively, so as to determine whether the person is wearing the reflective clothing; S4. If the person does not wear reflective clothing, an alarm signal is output.
[0006] Furthermore, step S1 specifically includes: Use deep learning-based human detection models to identify people in video streams; Using a multi-target tracking algorithm to keep tracking each of the detected persons; According to the pixel speed of each person and the rotation time t of the ball camera, the target position of each person after t time is predicted.
[0007] Furthermore, step S2 specifically includes: Determine in advance a number of the ball cameras associated with the main camera according to the installation position and coverage area of the main camera; Controlling each of the ball cameras to rotate to the predicted target position of the person; Performing secondary recognition of the person in the ball camera screen of each ball camera, and determining the ball camera that recognizes the person as the associated ball camera; The shooting parameters of each of the associated ball cameras are adjusted so that the person is located in the center of the ball camera screen, and the ball camera screen is enlarged.
[0008] Furthermore, step S3 specifically includes: In the center-enlarged ball camera image, a human body detection model based on deep learning is used to identify the body image of the person; The segformer semantic segmentation model is used to segment the upper body image of the person from the body image, and the segformer semantic segmentation model is used to segment the reflective clothing area image from the upper body image, so as to determine whether the person is wearing reflective clothing.
[0009] Furthermore, the step of segmenting the upper body image of the person from the body image using the segformer semantic segmentation model includes: Determining, based on the body image, whether a pixel aspect ratio of the person is less than a first preset value; If yes, it is determined that the person is in a non-sideways state, and the upper body image of the person is obtained by segmenting the body image using a segformer semantic segmentation model; If not, it is determined that the person is in a sideways state, and the result of the current frame is marked, and the result weight corresponding to the frame is correspondingly reduced when multiple frame results are subsequently fused.
[0010] Furthermore, the use of the segformer semantic segmentation model to segment the reflective clothing area image in the upper body image, thereby determining whether the person is wearing the reflective clothing includes: Segmenting the reflective clothing area image in the upper body image using the segformer semantic segmentation model, and calculating the ratio between the total number of pixels N of the reflective clothing area obtained by segmentation and the total number of pixels M of the upper body image of the person; If the ratio is greater than a second preset value, it is determined that the person in the frame is wearing reflective clothing; otherwise, it is determined that the person in the frame is not wearing reflective clothing; The percentage of the number of frames in which the person is determined to be wearing reflective clothing to the total number of all frames is calculated. If the percentage is higher than a third preset value, it is determined that the person is wearing reflective clothing; otherwise, it is determined that the person is not wearing reflective clothing.
[0011] Optionally, step S4 specifically includes: Summary of multiple people in the dome camera: When multiple people are found in the associated dome camera, an alarm signal is output if one person is found not wearing reflective clothing; Multi-ball machine result summary: 1) When the number of the associated dome cameras is 1, directly output the analysis results of the current associated dome cameras on each of the personnel; wherein the analysis results include the result of whether the personnel are wearing reflective clothing, and the corresponding alarm signal when the personnel are not wearing reflective clothing; 2) When the number of associated dome cameras is greater than 1: a) When the number of people in all the associated dome cameras is 1, the analysis result of the associated dome camera with the most analysis results is output; if the number of analysis results of the associated dome cameras is the same, the analysis result of the associated dome camera that can see the person from the non-side view is output; b) If the number of the persons in the associated dome cameras is greater than 1, if any dome camera generates an alarm, it is considered that one of the persons is not wearing a reflective vest.
[0012] The second aspect of the present invention provides a system for detecting reflective clothing for people in a large scene, comprising a main camera, a plurality of associated ball cameras, a processing module, and a storage module; the processing module is connected to the main camera, each of the ball cameras, and the storage module; The storage module is used to store executable computer program code; The main camera and the associated ball camera are used to capture images containing people and transmit them to the processing module; The processing module is used to execute the method as described in any of the preceding items by calling the executable computer program code in the storage module.
[0013] The third aspect of the present invention provides an electronic device, comprising: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the method as described in any of the preceding items.
[0014] A fourth aspect of the present invention provides a computer storage medium having a computer program stored thereon, and the computer program, when executed by a processor, executes any of the methods described above.
[0015] A fifth aspect of the present invention further provides a computer program product, comprising a computer program stored on a non-transitory computer-readable medium, wherein the computer program is executed by a processor to perform the method as described in any of the preceding items.
[0016] The beneficial effects of the present invention are at least: 1) It can detect the reflective clothing worn by small targets (people) in large scene images, avoiding inaccurate detection caused by the small proportion of the person image; 2) It breaks through the limitation of conventional deep learning-based object detection algorithms’ dependence on clothing types, reducing the problems of excessive sample requirements and difficulty in model updating caused by a wide variety of clothing; 3) It can effectively handle problems such as occlusion and overlap of people in large-scene videos, and the small proportion of people's side reflective clothing, reduce the rate of missed reports and false alarms, and improve the accuracy and reliability of overall detection; 4) It has strong adaptability and scalability, and can be used to detect reflective clothing of people in various large scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 It is a flow chart of a method for detecting reflective clothing of a person in a large scene disclosed in an embodiment of the present invention.
[0019] Figure 2 It is a structural schematic diagram of a system for detecting reflective clothing for people in a large scene disclosed in an embodiment of the present invention.
[0020] Figure 3 It is a structural schematic diagram of an electronic device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings, and "multiple" generally includes at least two.
[0023] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0024] It should be understood that although the terms first, second, third, etc. may be used to describe ... in the embodiments of the present application, these ... should not be limited to these terms. These terms are only used to distinguish .... For example, without departing from the scope of the embodiments of the present application, the first ... may also be referred to as the second ..., and similarly, the second ... may also be referred to as the first ....
[0025] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0026] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a product or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such a product or system. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the product or system including the elements.
[0027] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0028] like Figure 1 As shown, the embodiment of the present invention discloses a method for detecting reflective clothing of a person in a large scene, comprising the following steps: S1. Use a main camera to capture a video stream in real time, identify persons appearing in the video stream, and detect, track and predict each of the persons to obtain a predicted target position of each of the persons; S2, determining a number of associated ball cameras according to the target position, and controlling at least one of the associated ball cameras to perform centering and zooming processing on the person in the ball camera screen; S3, in the ball camera image after the center enlargement processing, using the segmentation model to segment the upper body and reflective clothing of the person respectively, so as to determine whether the person is wearing the reflective clothing; S4. If the person does not wear reflective clothing, an alarm signal is output.
[0029] Compared with the existing methods mentioned in the background technology, the above-mentioned scheme of the present invention respectively utilizes means such as personnel tracking, multi-ball camera linkage, image centering and magnification processing, and target segmentation, which can realize the detection of small targets (personnel) wearing reflective clothing in large scene images, avoiding the problem of inaccurate detection caused by the small proportion of personnel images; and can effectively deal with problems such as personnel occlusion, overlap, and a small proportion of personnel side reflective clothing in large scene videos, reduce the missed alarm and false alarm rates, and improve the accuracy and reliability of the overall detection.
[0030] It should be noted that the main camera in the present invention can be a preset camera for identifying personnel, or it can be any camera in the monitoring area, that is, it can be a gun camera or a ball camera, and the present invention does not make specific limitations on this.
[0031] Furthermore, step S1 specifically includes: Use deep learning-based human detection models to identify people in video streams; Using a multi-target tracking algorithm to keep tracking each of the detected persons; According to the pixel speed of each person and the rotation time t of the ball camera, the target position of each person after t time is predicted.
[0032] In this embodiment, this step is to ensure that the identified person can be in the camera screen as much as possible after the camera is triggered to zoom in in the subsequent steps, and to reserve sufficient analysis time for the reflective clothing to be analyzed after the camera zooms in. Among them, the human body detection model can use, for example, YOLO11, and the camera rotation time t can be a priori knowledge, which is estimated in the early stage based on information such as different camera positions.
[0033] An improved solution of this embodiment is provided: the method of using a multi-target tracking algorithm to continuously track each of the detected persons includes: Continuously tracking each of the personnel within a preset time period, and counting the fluctuation range of the number of the personnel within the preset time period, and determining the extension time period according to the fluctuation range of the number; wherein the extension time period is positively correlated with the fluctuation range of the number; Continue to follow up on each of the persons during the extended period.
[0034] When there are many people in the video stream, overlap, occlusion, etc. will cause a deviation in the prediction of the target position of the person being tracked after time t, and the deviation will increase with the number of people in the video stream. In this regard, the present invention is set to first continuously track each person within a preset time length (for example, 5s), and count the fluctuation range of the number of people within the preset time length (overlap, occlusion, etc. will cause the number of people in each frame of the video image to fluctuate, decrease or increase), and determine the extended time length according to the positive correlation relationship based on the fluctuation range of the number (absolute value), that is, when the fluctuation range of the number is larger, the longer the extended time length is set, that is, the prediction accuracy of the target position of the person is ensured by tracking for a longer time.
[0035] Furthermore, step S2 specifically includes: Determine in advance a number of the ball cameras associated with the main camera according to the installation position and coverage area of the main camera; Controlling each of the ball cameras to rotate to the predicted target position of the person; Performing secondary recognition of the person in the ball camera screen of each ball camera, and determining the ball camera that recognizes the person as the associated ball camera; The shooting parameters of each of the associated ball cameras are adjusted so that the person is located in the center of the ball camera screen, and the ball camera screen is enlarged.
[0036] In this embodiment, some cameras can be associated, i.e., offline bound, based on factors such as proximity of the installation location and coverage area, as well as whether they are located in an area where people can walk, to facilitate subsequent triggering of the linkage between the main camera and the associated dome camera, ensuring that the dome camera can capture people from multiple angles.
[0037] Since the rotation and zooming of the ball camera requires a certain amount of time, this time (such as t seconds) is related to the rotation mechanism of the ball camera itself and the degree of rotation and zooming required. Therefore, in order to ensure that the target is centered in the ball camera screen as much as possible after rotation, the ball camera needs to be rotated to the predicted target position of the person after t seconds. Since the rotation speed of the ball camera is usually fast and the zooming is slow, t here is set to be related to the zooming (Z). Usually, the (ti, Zi) table is set according to the actual situation. In actual use, the corresponding t value is obtained by table lookup. Therefore, the position of the corresponding person after the rotation and zooming of the ball camera is calculated according to the prediction method in step S1.
[0038] In addition, it is necessary to further explain: 1) In order to ensure that the reflective vests of personnel can be clearly seen from multiple angles, usually at least 2 dome cameras are linked (to view personnel from mutually perpendicular angles). More dome cameras (such as 4) can also be called according to the actual situation on site to view personnel from multiple angles such as front, back, left, and right; 2) The multi-angle viewing of the ball camera can solve the problem of one or more people being blocked at certain angles, thereby solving the problem of underreporting caused by conventional methods; 3) The multi-angle viewing of the ball camera can also solve the problem of false alarms caused by people appearing sideways in the video at a single angle (when people are sideways, the reflective clothing occupies a very small proportion of the person's pixels, resulting in various models on the market may not be able to accurately detect them); 4) The multi-angle viewing of the ball camera can avoid the overlap of people in the video at some angles (the probability of people overlapping at multiple angles will be greatly reduced). This overlap will cause the conventional reflective clothing detection algorithm on the market to detect that the person area is carrying part of the body (including reflective clothing) of other people, resulting in false alarms or missed alarms.
[0039] Furthermore, step S3 specifically includes: In the center-enlarged ball camera image, a human body detection model based on deep learning is used to identify the body image of the person; The segformer semantic segmentation model is used to segment the upper body image of the person from the body image, and the segformer semantic segmentation model is used to segment the reflective clothing area image from the upper body image, so as to determine whether the person is wearing reflective clothing.
[0040] In this embodiment, after the identified person is centered and enlarged, it is necessary to identify the person in the centered and enlarged ball camera image again (YOLO11 can also be used) to determine the body image of the person in the centered and enlarged ball camera image. Then, the preset segformer semantic segmentation model is used to segment the upper body image of the person, and the reflective clothing area image is segmented from the upper body image.
[0041] Furthermore, the step of segmenting the upper body image of the person from the body image using the segformer semantic segmentation model includes: Determining, based on the body image, whether a pixel aspect ratio of the person is less than a first preset value; If yes, it is determined that the person is in a non-sideways state, and the upper body image of the person is obtained by segmenting the body image using a segformer semantic segmentation model; If not, it is determined that the person is in a sideways state, and the result of the current frame is marked, and the result weight corresponding to the frame is correspondingly reduced when multiple frame results are subsequently fused.
[0042] In this embodiment, in order to prevent the model from being inaccurate due to the sideways position of a person, it is necessary to determine whether the person is sideways when the upper body of the person is segmented using the segformer semantic segmentation model. For example, it is determined whether the pixel aspect ratio of the person is less than 6 (the first preset value). If so, it indicates that the person is not sideways, otherwise, it indicates that the person is sideways.
[0043] If the person is judged to be in a frontal state, the segmentation model can be directly used to better segment the upper body of the person, thereby solving the problem of false alarms in subsequent reflective clothing detection caused by the introduction of other people's bodies and clothes into the boudingbox of the upper body of the person detected by the conventional detection model when the person overlaps. If the person is judged to be in a sideways state, the result of the current frame (i.e., whether the reflective clothing is worn) needs to be annotated, and the results of multiple frames need to be fused later. When fusion occurs, the result weights corresponding to each frame are different, and the fusion weight corresponding to the frame in the sideways state is correspondingly reduced based on the above annotations.
[0044] Furthermore, the use of the segformer semantic segmentation model to segment the reflective clothing area image in the upper body image, thereby determining whether the person is wearing the reflective clothing includes: Segmenting the reflective clothing area image in the upper body image using the segformer semantic segmentation model, and calculating the ratio between the total number of pixels N of the reflective clothing area obtained by segmentation and the total number of pixels M of the upper body image of the person; If the ratio is greater than a second preset value, it is determined that the person in the frame is wearing reflective clothing; otherwise, it is determined that the person in the frame is not wearing reflective clothing; The percentage of the number of frames in which the person is determined to be wearing reflective clothing to the total number of all frames is calculated. If the percentage is higher than a third preset value, it is determined that the person is wearing reflective clothing; otherwise, it is determined that the person is not wearing reflective clothing.
[0045] In this embodiment, by segmenting the ratio of the total number of pixels N in the reflective clothing area to the total number of pixels M in the upper body image of the person, if it is greater than 0.3 (the second preset value), it means that the reflective clothing is worn, otherwise it is considered that it is not worn, and continuous judgment is made over multiple frames, for example, more than 60% (the third preset value) of the judgment results in 50 frames are judged as wearing reflective clothing, so as to comprehensively judge whether the person in the current ball camera viewing angle is wearing reflective clothing.
[0046] The main reasons for choosing the segformer semantic segmentation model here are: Efficiency: Segformer is relatively efficient in terms of computing resources and model parameters. It reduces the number of model parameters and computational complexity through a simple decoder design and a hierarchical Transformer encoder. Compared with some traditional complex CNN-based segmentation models, it can maintain high segmentation accuracy while having a faster inference speed, making it suitable for applications in scenarios with high real-time requirements.
[0047] Multi-scale feature utilization: Thanks to its hierarchical encoder structure, Segformer can make good use of multi-scale features. In the semantic segmentation task, objects and scene details of different scales can be effectively captured and processed. This multi-scale feature fusion method helps to improve the model's segmentation accuracy for objects of different sizes.
[0048] Adaptability and generalization: Due to the characteristics of the Transformer architecture itself, Segformer has good adaptability to different types of image data and segmentation tasks. It does not rely on fixed convolution kernels and local receptive fields like some CNN models trained based on manual features or specific data sets. The Transformer's self-attention mechanism can automatically learn long-distance dependencies between features, making the model more generalizable when facing different image distributions and task requirements.
[0049] Optionally, step S4 specifically includes: Summary of multiple people in the dome camera: When multiple people are found in the associated dome camera, an alarm signal is output if one person is found not wearing reflective clothing; Multi-ball machine result summary: 1) When the number of the associated dome cameras is 1, directly output the analysis results of the current associated dome cameras on each of the personnel; wherein the analysis results include the result of whether the personnel are wearing reflective clothing, and the corresponding alarm signal when the personnel are not wearing reflective clothing; 2) When the number of associated dome cameras is greater than 1: a) When the number of people in all the associated dome cameras is 1, the analysis result of the associated dome camera with the most analysis results is output; if the number of analysis results of the associated dome cameras is the same, the analysis result of the associated dome camera that can see the non-side view of the person is output (when there are 2 or more linked dome cameras, when selecting the dome camera angle, at least two dome cameras are substantially perpendicular to each other, that is, there must be at least one dome camera that can see the front or back of the person); b) If the number of persons in the associated dome cameras is greater than 1 (in this case, there may be problems such as obstruction and overlap), if any dome camera generates an alarm, it is considered that one of the persons is not wearing a reflective vest.
[0050] like Figure 2 As shown, the embodiment of the present invention discloses a system for detecting reflective clothing of personnel in a large scene, including a main camera, a plurality of associated ball cameras, a processing module, and a storage module; the processing module is connected with the main camera, each of the ball cameras, and the storage module; The storage module is used to store executable computer program code; The main camera and the associated ball camera are used to capture images containing people and transmit them to the processing module; The processing module is used to execute the method as described in any of the preceding items by calling the executable computer program code in the storage module.
[0051] The specific functions of the reflective clothing detection system for personnel in a large scene in this embodiment refer to the above embodiment. Since the system of this embodiment adopts all the technical solutions of the above embodiment, it at least has all the beneficial effects brought by the technical solutions of the above embodiment, which will not be described one by one here.
[0052] like Figure 3 As shown, an embodiment of the present invention discloses an electronic device, comprising: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the method described in the above embodiment.
[0053] The embodiment of the present invention further discloses a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the above embodiment is executed.
[0054] An embodiment of the present invention further discloses a computer program product, including a computer program stored on a non-transitory computer-readable medium, wherein the computer program is executed by a processor to perform the method described in the above embodiment.
[0055] The device / system according to an embodiment of the present disclosure may include a processor, a memory for storing program data and executing the program data, a permanent memory such as a disk drive, a communication port for processing communication with an external device, and a user interface device, etc. The method is implemented as a software module or can be stored on a computer-readable recording medium as a computer-readable code or program command that can be executed by a processor. Examples of computer-readable recording media may include magnetic storage media (e.g., read-only memory (ROM), random access memory (RAM), floppy disk, hard disk, etc.), optical reading media (e.g., CD-ROM, digital versatile disk (DVD), etc.), etc. The computer-readable recording medium may be distributed in a computer system connected in a network, and the computer-readable code may be stored and executed in a distributed manner. The medium may be computer-readable, stored in a memory and executed by a processor.
[0056] The embodiments of the present disclosure may be indicated as function block components and various processing operations. Function blocks may be implemented as various numbers of hardware and / or software components that perform specific functions. For example, the embodiments of the present disclosure may implement direct circuit components that can perform various functions under the control of one or more microprocessors or other control devices, such as memory, processing circuits, logic circuits, lookup tables, etc. The components of the present disclosure may be implemented by software programming or software components. Similarly, the embodiments of the present disclosure may include various algorithms implemented by a combination of data structures, processes, routines, or other programming components, and may be implemented by programming or scripting languages (such as C, C++, Java, assemblers, etc.). Functional aspects may be implemented by algorithms executed by one or more processors. In addition, the embodiments of the present disclosure may implement related technologies for electronic environment settings, signal processing, and / or data processing. Terms such as "mechanism", "element", "unit", etc. may be used extensively and are not limited to mechanical and physical components. These terms may represent a series of software routines associated with a processor, etc.
[0057] Specific embodiments are described in the present disclosure as examples, and the scope of the embodiments is not limited thereto.
[0058] Although the embodiments of the present disclosure have been described, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims. Therefore, the above embodiments of the present disclosure should be interpreted as examples and do not limit the embodiments in all aspects. For example, each component described as a single unit may be performed in a distributed manner, and similarly, components described as distributed may be performed in a combined manner.
[0059] All examples or exemplary terms (for example, etc.) used in the embodiments of the present disclosure are for the purpose of describing the embodiments of the present disclosure, but are not intended to limit the scope of the embodiments of the present disclosure.
[0060] Furthermore, unless explicitly stated otherwise, expressions such as “essential,” “important,” etc., associated with certain components may not indicate that the components are absolutely required.
[0061] It will be appreciated by those skilled in the art that the embodiments of the present disclosure may be implemented in modified forms without departing from the spirit and scope of the present disclosure.
[0062] Since the present disclosure allows various changes to the embodiments of the present disclosure, the present disclosure is not limited to specific embodiments, and it will be understood that all changes, equivalents and substitutes that do not depart from the spirit and technical scope of the present disclosure are included in the present disclosure. Therefore, the embodiments of the present disclosure described herein should be understood as examples in all aspects and should not be interpreted as limitations.
[0063] In addition, terms such as "unit", "module", etc. represent a unit that can be implemented as hardware or software or a combination of hardware and software to process at least one function or operation. "Units" and "modules" can be stored in a storage medium to be addressed, and can be implemented as a program that can be executed by a processor. For example, "units" and "modules" can refer to components such as software components, object-oriented software components, class components, and task components, and can include processes, functions, properties, procedures, subroutines, program code segments, drivers, firmware, microcodes, circuits, data, databases, data structures, tables, arrays, or variables.
[0064] In the present disclosure, the expression "A may include one of a1, a2, and a3" may broadly indicate that examples that may be included in element A include a1, a2, or a3. The expression should not be interpreted as being limited to the meaning that examples included in element A must be limited to a1, a2, and a3. Therefore, as examples included in element A, it should not be interpreted as excluding elements other than a1, a2, and a3. In addition, the expression indicates that element A may include a1, a2, or a3. The expression does not mean that the elements included in element A must be selected from a specific set of elements. That is, the expression should not be restrictively understood as indicating that a1, a2, or a3 that must be selected from the set including a1, a2, and a3 is included in element A.
[0065] Furthermore, in the present disclosure, the expression “at least one of a1, a2, and / or a3” means one of “a1,” “a2,” “a3,” “a1 and a2,” “a1 and a3,” “a2 and a3,” and “a1, a2, and a3.” Therefore, it should be noted that the expression “at least one of a1, a2, and / or a3” should not be interpreted as “at least one of a1,” “at least one of a2,” and “at least one of a3,” unless explicitly described as “at least one of a1, at least one of a2, and at least one of a3.”
Claims
1. A method for detecting reflective clothing of people in a large scene, characterized in that: The steps include: S1. Use a main camera to capture a video stream in real time, identify persons appearing in the video stream, and detect, track and predict each of the persons to obtain a predicted target position of each of the persons; S2, determining a number of associated ball cameras according to the target position, and controlling at least one of the associated ball cameras to perform centering and zooming processing on the person in the ball camera screen; S3, in the ball camera image after the center enlargement processing, using the segmentation model to segment the upper body and reflective clothing of the person respectively, so as to determine whether the person is wearing the reflective clothing; S4. If the person does not wear reflective clothing, an alarm signal is output.
2. According to the method for detecting reflective clothing of people in a large scene in claim 1, it is characterized by: Step S1 specifically includes: Use deep learning-based human detection models to identify people in video streams; Using a multi-target tracking algorithm to keep tracking each of the detected persons; According to the pixel speed of each person and the rotation time t of the ball camera, the target position of each person after t time is predicted.
3. According to the method for detecting reflective clothing of people in a large scene in claim 2, it is characterized by: Step S2 specifically includes: Determine in advance a number of the ball cameras associated with the main camera according to the installation position and coverage area of the main camera; Controlling each of the ball cameras to rotate to the predicted target position of the person; Performing secondary recognition of the person in the ball camera screen of each ball camera, and determining the ball camera that recognizes the person as the associated ball camera; The shooting parameters of each of the associated ball cameras are adjusted so that the person is located in the center of the ball camera screen, and the ball camera screen is enlarged.
4. According to the method for detecting reflective clothing of people in a large scene in claim 3, it is characterized by: Step S3 specifically includes: In the center-enlarged ball camera image, a human body detection model based on deep learning is used to identify the body image of the person; The segformer semantic segmentation model is used to segment the upper body image of the person from the body image, and the segformer semantic segmentation model is used to segment the reflective clothing area image from the upper body image, so as to determine whether the person is wearing reflective clothing.
5. The method for detecting reflective clothing of people in a large scene according to claim 4, characterized in that: The step of using the segformer semantic segmentation model to segment the upper body image of the person from the body image includes: Determining, based on the body image, whether a pixel aspect ratio of the person is less than a first preset value; If yes, it is determined that the person is in a non-sideways state, and the upper body image of the person is obtained by segmenting the body image using a segformer semantic segmentation model; If not, it is determined that the person is in a sideways state, and the result of the current frame is marked, and the result weight corresponding to the frame is correspondingly reduced when multiple frame results are subsequently fused.
6. The method for detecting reflective clothing of people in a large scene according to claim 4, characterized in that: The step of using the segformer semantic segmentation model to segment the reflective clothing area image in the upper body image, thereby determining whether the person is wearing the reflective clothing, includes: Segmenting the reflective clothing area image in the upper body image using the segformer semantic segmentation model, and calculating the ratio between the total number of pixels N of the reflective clothing area obtained by segmentation and the total number of pixels M of the upper body image of the person; If the ratio is greater than a second preset value, it is determined that the person in the frame is wearing reflective clothing; otherwise, it is determined that the person in the frame is not wearing reflective clothing; The percentage of the number of frames in which the person is determined to be wearing reflective clothing to the total number of all frames is calculated. If the percentage is higher than a third preset value, it is determined that the person is wearing reflective clothing; otherwise, it is determined that the person is not wearing reflective clothing.
7. The method for detecting reflective clothing of people in a large scene according to claim 1, characterized in that: Step S4 specifically includes: Summary of multiple people in the dome camera: When multiple people are found in the associated dome camera, an alarm signal is output if one person is found not wearing reflective clothing; Multi-ball machine result summary: 1) When the number of the associated dome cameras is 1, directly output the analysis results of the current associated dome cameras on each of the personnel; wherein the analysis results include the result of whether the personnel are wearing reflective clothing, and the corresponding alarm signal when the personnel are not wearing reflective clothing; 2) When the number of associated dome cameras is greater than 1: a) When the number of people in all the associated dome cameras is 1, the analysis result of the associated dome camera with the most analysis results is output; if the number of analysis results of the associated dome cameras is the same, the analysis result of the associated dome camera that can see the person from the non-side view is output; b) If the number of the persons in the associated dome cameras is greater than 1, if any dome camera generates an alarm, it is considered that one of the persons is not wearing a reflective vest.
8. A system for detecting reflective clothing for people in a large scene, comprising a main camera, several associated ball cameras, a processing module, and a storage module; the processing module is connected to the main camera, each of the ball cameras, and the storage module; The storage module is used to store executable computer program code; The main camera and the associated ball camera are used to capture images containing people and transmit them to the processing module; Features: The processing module is used to execute the method according to any one of claims 1 to 7 by calling the executable computer program code in the storage module, so as to realize the superimposed display of the cargo content in the video image.
9. An electronic device, comprising: A memory storing executable program code; A processor coupled to the memory; characterized in that: the processor calls the executable program code stored in the memory to execute the method according to any one of claims 1-7.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Safety helmet wearing detection and tracking method based on improved YOLOv3
CN110852283A
Indoor Internet of Things video tracking method and system
CN111311649A
Reflective rope wearing detection method, device and equipment and storage medium
CN113903055A
Method and device for detecting dangerous behaviors in industrial scene and medium
CN117912101A
Video analytics for industrial floor setting
US20240281954A1