Method and device for detecting position of judge in standard court, method and device for detecting whether judge wears gown or not, and medium

Through the improved yolov8s and Resnet50 models combined with the Deepsort algorithm, the accuracy and robustness of judge robes recognition are solved, real-time monitoring and accurate identification of judge dress.

CN120339684AActive Publication Date: 2025-07-18BEIJING DONGFANG GUOZHENG INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510339715.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

Traditional methods rely on manual judgment on whether judges are wearing robes in standardized manner, which is inefficient and subjective. The existing image processing technology lacks accuracy and robustness in identifying judges wearing robes in complex court environments.

Method used

The personnel detection model based on yolov8s is used to combine the Resnet50 model and the Deepsort algorithm, and the fusion of wavelet convolution and frequency domain perceptual feature is used to identify the judge's position and judge the wearing of the robe, and use HSV histogram information for real-time monitoring.

Benefits of technology

It realizes rapid and accurate identification of judges' dressing situations, maintains stable identification performance in complex court environments, and improves the accuracy and robustness of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339684A_ABST
    Figure CN120339684A_ABST
Patent Text Reader

Abstract

The invention provides a method for detecting the position of a judge in a standard court, a method and device for detecting whether the judge wears a gown, and a medium, and belongs to the technical field of image processing. In a court environment, the position of a judge is relatively fixed, so that the position of the judge is firstly identified by a people detection model based on the yov8s; and after identification, further utilizing an image classification technology to divide the detected judge images into two types, namely wearing a gown and not wearing the gown. In consideration of the shielding problem possibly occurring in the actual situation, the detected judge needs to be continuously tracked, so that continuous monitoring is ensured. The process can depend on a target tracking technology, the position of the judge wearing the gown is tracked and judged in real time according to the extracted coordinates of the judge, whether the judge regularly wears the gown or not is judged by counting histogram information of three HSV channels of the position area where the judge is located, and continuous visibility and correct classification of the judge are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting the position of a judge in a courtroom and a method, device, and medium for detecting whether a judge is wearing a robe, belonging to the technical field of image processing. Background Art

[0002] In the modern judicial system, the dress code of judges in courtrooms is an important manifestation of respecting and maintaining the dignity of the law. However, for the judgment of whether judges are dressed in robes in courtrooms, traditional methods mainly rely on manual observation and judgment, which is not only inefficient but also may be affected by human factors, resulting in certain subjectivity and inaccuracy in the judgment results. With the development of computer vision and deep learning technologies, image processing technologies have been widely applied in various scenarios, including face recognition, object detection, and behavior analysis. However, in courtroom scenarios, especially for the automatic recognition task of whether judges are wearing robes, although there have been some related studies and attempts, there are still some challenges and problems. For example, the lighting conditions in the courtroom environment are complex, the positions of judges may change, the colors and textures of robes may vary due to individual differences, and there may even be occlusions, etc., which all pose challenges to accurate recognition.

[0003] Facing these challenges, traditional object detection technologies such as template matching and feature classifiers encounter limitations when dealing with complex courtroom environments. In contrast, modern object detection technologies based on deep learning, such as convolutional neural networks and related models (R-CNN series, SSD, YOLO series, etc.), automatically extract features by learning a large amount of labeled data, not only perform well in object classification and localization, but also can more effectively handle problems such as light changes, occlusions, and position changes. These technologies improve the accuracy and robustness of recognition by deeply learning the visual features of robes, including colors, textures, and shapes, providing a more efficient and reliable technical means for automatically recognizing the dressing standards of judges in courtroom scenarios. With the development of emerging technologies such as attention mechanism networks, end-to-end object detection systems, and generative adversarial networks (GANs), the application of deep learning in courtroom environments is expected to be further innovated and optimized, improving the efficiency and fairness of the judicial system.

[0004] Meanwhile, in the aspect of image classification, traditional machine learning methods rely on manually extracting features, such as SIFT or HOG, and combining classifiers such as support vector machines (SVMs) or random forests for classification. Although these methods perform well on small-scale datasets, they may encounter bottlenecks in large-scale and complex scenarios. Deep learning techniques such as convolutional neural networks (CNNs) have significantly improved the image classification performance on large-scale datasets by automatically learning feature representations. At the same time, through innovative network structures and training techniques, they have effectively solved the challenges in the training of deep networks, making the application of deep learning in image classification tasks more extensive and in-depth. Summary of the Invention

[0005] The object of the present invention is to provide a method and device, and medium for detecting the position of a judge and whether a judge is wearing a robe in a standardized courtroom, which improves the accuracy and robustness of recognition.

[0006] To achieve the above object, the present invention is realized through the following technical solutions: A method for detecting the position of a judge in a standardized courtroom, comprising: Collecting image samples of a standardized courtroom with and without a judge, constructing a personnel detection dataset, and annotating the position of the judge; Constructing a personnel detection model based on yolov8s; the personnel detection model includes a backbone network, a frequency-domain perception feature fusion module, and a detection head. The outputs of the second convolution module, the third convolution module, and the fourth convolution module of the backbone network are all input into the feature extraction module after passing through the wavelet convolution module; the frequency-domain perception feature fusion module sequentially includes a first frequency-domain perception fusion module, a second frequency-domain perception fusion module, a third frequency-domain perception fusion module, a first feature extraction module, a second feature extraction module, and a third feature extraction module; Training the personnel detection model based on yolov8s using the dataset; Collecting an image of the judge's bench position in the courtroom and inputting it into the trained personnel detection model to obtain the position of the judge.

[0007] Preferably, the backbone network of the personnel detection model based on yolov8s sequentially includes a first convolution module, a second convolution module, a first wavelet convolution, a first feature extraction module, a third convolution module, a second wavelet convolution, a second feature extraction module, a fourth convolution module, a third wavelet convolution, a third feature extraction module, a fifth convolution module, a fourth feature extraction module, and an SPPF module.

[0008] Preferably, the processing method of the feature map of the personnel detection model is as follows: Input the image into the backbone network; the feature map output by the SPPF module of the backbone network and the feature map output by the fourth feature extraction module are input into the first frequency-domain perception fusion module of the frequency-domain perception feature fusion module, and the obtained feature maps a1, a2, and a3 are output. Add feature map a1 and feature map a2 to obtain feature map a4. Among them, feature map a3 is the mask of the filter, and feature map a1 and feature map a2 are high-resolution feature maps after frequency fusion; The feature map output by the second feature extraction module of the backbone network and feature map a4 are input into the second frequency-domain perception fusion module of the frequency-domain perception feature fusion module, and the obtained feature maps b1, b2, and b3 are output. Add feature map b1 and feature map b2 to obtain feature map b4. Among them, feature map b3 is the mask of the filter, and feature map b1 and feature map b2 are high-resolution feature maps after frequency fusion; The feature map output by the first feature extraction module of the backbone network and feature map b4 are input into the third frequency-domain perception fusion module of the frequency-domain perception feature fusion module, and the obtained feature maps c1, c2, and c3 are output. Add feature map c1 and feature map c2 to obtain feature map c4. Among them, feature map c3 is the mask of the filter, and feature map c1 and feature map c2 are high-resolution feature maps after frequency fusion; Input feature map c4 into the first feature extraction module of the frequency-domain perception feature fusion module, and the obtained feature map F1 is output; Concatenate feature map F1 and feature map b4 in the channel dimension, and input them into the second feature extraction module of the frequency-domain perception feature fusion module, and the obtained feature map F2 is output; Concatenate feature map F2 and feature map a3 in the channel dimension, and input them into the third feature extraction module of the frequency-domain perception feature fusion module, and the obtained feature map F3 is output; Input feature map F1, feature map F2, and feature map F3 into the detection head to obtain the detection result.

[0009] A method for detecting whether a judge wears a robe in a standardized courtroom includes: Collect image samples of a judge wearing a robe and not wearing a robe in a standardized courtroom, cut the pictures with the minimum bounding rectangle of the judge's body contour and save them. Divide the cut pictures into two categories: wearing a robe and not wearing a robe, and construct a classification data set; Use the classification data set to train the resnet50 model; Collect real-time images of the standardized courtroom scene. The person detection model of the method for detecting the position of a judge in the standardized courtroom detects the position of the judge and extracts the coordinates of the judge in the image; Cut the images of all detected regions where judges are present using the minimum bounding rectangle of the judge's body contour, and input the cut images into the trained Resnet50 model for classification to determine whether all judges are wearing judicial robes; Track the positions of the judges wearing judicial robes in real time based on the extracted judge coordinates, and determine whether the judges are wearing judicial robes properly by statistically analyzing the histogram information of the HSV channels in the regions where the judges are located.

[0010] Preferably, track the positions of the judges wearing judicial robes in real time based on the extracted judge coordinates, and determine whether the judges are wearing judicial robes properly by statistically analyzing the histogram information of the HSV channels in the regions where the judges are located. The specific method is as follows: Input the extracted judge coordinates into the Deepsort algorithm for tracking, and assign numbers to the detected judges; For the judges wearing judicial robes, according to the positions tracked by the Deepsort algorithm, collect the position regions every seconds during the court trial in real time, and statistically analyze the histogram information of the HSV channels in these regions; Calculate the Bhattacharyya distance between the current moment and the histogram of the corresponding channels seconds ago. After obtaining the Bhattacharyya distances of the three channels, calculate their average value. If the average value is less than the set threshold, it is determined that the judge is wearing a judicial robe; if it is greater than the preset threshold, it means that the judge is not wearing a judicial robe.

[0011] Preferably, calculate the Bhattacharyya distance between the current moment and the histogram of the corresponding channels seconds ago. The specific method is as follows: Obtain the histogram statistical information of the HSV channels in the cut image regions with the same number as the current moment and seconds ago; Calculate the Bhattacharyya distance between the histograms of the corresponding channels at the two moments. First, calculate the Bhattacharyya coefficient, and the formula is as follows: , where, represents the pixel-level range, represents the histogram of the H, S, or V channel seconds ago, represents the histogram of the H, S, or V channel at the current moment, represents the Bhattacharyya coefficient; Calculate the Bhattacharyya distances of all channels according to the Bhattacharyya distance formula. The Bhattacharyya distance formula is as follows: , where, represents the Bhattacharyya distance between the histogram of the H, S, or V channel seconds ago and the histogram of this channel at the current moment.

[0012] Preferably, the takes any integer number of seconds within the range of [1, 10].

[0013] Preferably, the method for selecting the set threshold is as follows: Collect image samples of all judges wearing and not wearing judicial robes within a continuous period of time. Take the image samples of a judge wearing and not wearing a judicial robe with the same number as a pair of samples. Respectively extract the histogram information of the HSV three channels of the area where the judge is located; Calculate the Bhattacharyya distance between the HSV three-channel histograms of each pair of samples, take the average value of the Bhattacharyya distances of all pairs of samples, and multiply by an adjustment coefficient as the set threshold; the adjustment coefficient α ranges from (0, 1).

[0014] A detection device for regulating whether a judge wears a judicial robe in a courtroom, including a processor and a memory storing program instructions. The processor is configured to execute the method for detecting whether a judge wears a judicial robe in a regulated courtroom when running the program instructions.

[0015] A computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for detecting whether a judge wears a judicial robe in a regulated courtroom.

[0016] The advantages of the present invention are as follows: The present invention can monitor the dressing situation of judges in the courtroom in real time, and use advanced deep learning models and object tracking algorithms to achieve fast and accurate identification of whether a judge wears a judicial robe. By improving the Yolov8s model using the wavelet convolution (WTConv) and FreqFusion methods, the present invention can more effectively handle complex situations such as light changes, angle deviations, and partial occlusions. This enables the system to maintain stable recognition performance in various courtroom environments, improving the accuracy and robustness of recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.

[0018] Figure 1 It is a schematic flowchart of the method for detecting the position of a judge in a regulated courtroom of the present invention.

[0019] Figure 2 It is a schematic flowchart of the method for detecting whether a judge wears a judicial robe in a regulated courtroom of the present invention.

[0020] Figure 3 It is a schematic diagram of the structure of the person detection model.

[0021] Figure 4Schematic diagram of the composition of the frequency-domain perception feature fusion module (FreqFusionNeck).

[0022] Figure 5 Schematic diagram of the overall process for standardizing the detection method of whether a judge wears a robe in court.

[0023] Figure 6 Schematic diagram of the image acquisition method. Specific implementation manners

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0025] Embodiment 1 In a court environment, the position of the judge is relatively fixed. Therefore, based on object detection technology, the position of the judge can be identified first. After identification, we can further use image classification technology to classify the detected judge images into two categories: wearing a robe and not wearing a robe. Considering the possible occlusion problems in actual situations, for the detected judge, we need to continuously track to ensure continuous monitoring. This process can rely on object tracking technology to ensure the continuous visibility and correct classification of the judge.

[0026] As Figure 1 shown, a method for detecting the position of a judge in a standardized court includes: S1: Collect image samples of the presence and absence of a judge in a standardized court, construct a personnel detection data set, and mark the position of the judge.

[0027] S2: Construct a personnel detection model based on yolov8s; the personnel detection model includes a backbone network, a frequency-domain perception feature fusion module, and a detection head. The outputs of the second convolution module, the third convolution module, and the fourth convolution module of the backbone network are all input into the feature extraction module after passing through the wavelet convolution module; the frequency-domain perception feature fusion module sequentially includes a first frequency-domain perception fusion module, a second frequency-domain perception fusion module, a third frequency-domain perception fusion module, a first feature extraction module, a second feature extraction module, and a third feature extraction module.

[0028] S3: Use the data set to train the personnel detection model based on yolov8s.

[0029] S4: Collect the image of the judge's bench position in the court and input it into the trained personnel detection model to obtain the position of the judge.

[0030] As a refinement of the above embodiment, samples of a judge wearing and not wearing a robe properly in court are collected through a camera installed in the court. For the personnel detection dataset, the position of the judge in the picture is manually annotated. When manually annotating the position of the judge in the picture, the circumscribed rectangle of the judge's body contour is used as the annotation standard, and the rectangle is accurately drawn to closely fit the judge's body area. At the same time, a class label is assigned to the rectangle to clearly define the judge object in the picture.

[0031] The personnel detection dataset should ensure that the number of the two types of samples in the dataset is as equal as possible. At the same time, the pictures in the dataset cover the judge images under various lighting conditions such as direct strong light, weak light scattering, side light, backlight, top light, front light, and natural and artificial mixed light. The images also include the judge's images presented from different angles such as the front, side, back, and oblique side, as well as the images under local occlusion conditions such as hand occlusion (such as raising the hand, flipping through documents), document occlusion (placing a large number of documents in front), microphone occlusion (due to microphone equipment), and other people occlusion (personnel interaction at the trial scene), so as to ensure the diversity and comprehensiveness of the data, enabling the model to accurately detect whether the judge is wearing a robe in various complex situations.

[0032] As a refinement of the above embodiment, as Figure 3 shown, the backbone network of the personnel detection model based on yolov8s sequentially includes a first convolution module (Conv), a second convolution module, a first wavelet convolution (WTConv), a first feature extraction module (CSPLayer), a third convolution module, a second wavelet convolution, a second feature extraction module, a fourth convolution module, a third wavelet convolution, a third feature extraction module, a fifth convolution module, a fourth feature extraction module, and an SPPF module.

[0033] The first convolution module, the second convolution module, the third convolution module, the fourth convolution module, and the fifth convolution module are all sequentially composed of a convolution layer, a BatchNormalization layer, and a SiLU activation function; the first feature extraction module, the second feature extraction module, the third feature extraction module, and the fourth feature extraction module are all composed of a convolution layer, a BatchNormalization layer, and a SiLU activation function; the first wavelet convolution, the second wavelet convolution, and the third wavelet convolution are all composed of a wavelet transform and a convolution layer.

[0034] Specifically, the feature processing process of the backbone network is as follows: The image is obtained and input into the first convolution module of the Backbone backbone network, and a feature map is output ; The feature map is input into the second convolution module of the Backbone backbone network, and a feature map is output ; Input the feature map into the first wavelet convolution module of the Backbone main network, and output to obtain the feature map ; Input the feature map into the first feature extraction module of the Backbone main network, and output to obtain the feature map Input the feature map into the third convolution module of the Backbone main network, and output to obtain the feature map ; Input the feature map into the second wavelet convolution module of the Backbone main network, and output to obtain the feature map ; Input the feature map into the second feature extraction module of the Backbone main network, and output to obtain the feature map ; Input the feature map into the fourth convolution module of the Backbone main network, and output to obtain the feature map ; Input the feature map into the third wavelet convolution module of the Backbone main network, and output to obtain the feature map ; Input the feature map into the fourth feature extraction module of the Backbone main network, and output to obtain the feature map ; Input the feature map into the sixth convolution module of the Backbone main network, and output to obtain the feature map ; Input the feature map into the fifth feature extraction module of the Backbone main network, and output to obtain the feature map ; Input the feature map into the SPPF module of the Backbone main network, and output to obtain the feature map .

[0035] As a refinement of the above embodiment, such as Figure 4As shown, the first frequency-domain perception fusion module (FreqFusion), the second frequency-domain perception fusion module, and the third frequency-domain perception fusion module of the frequency-domain perception feature fusion module are all composed of a convolutional layer, upsampling and downsampling operations, a feature resampling module (LocalSimGuidedSampler), Hamming window generation and application, feature compression, and feature normalization and initialization operations; the first feature extraction module, the second feature extraction module, and the third feature extraction module of the frequency-domain perception feature fusion module are all composed of a convolutional layer, a BatchNormalization layer, and a SiLU activation function.

[0036] Specifically, the feature processing flow of the frequency-domain perception feature fusion module is as follows: The SPPF module of the backbone network outputs a feature map and the output feature map of the fourth feature extraction module are input into the first frequency-domain perception fusion module of the frequency-domain perception feature fusion module, and feature maps a1, a2, and a3 are output. Feature map a1 and feature map a2 are added to obtain feature map a4. Among them, feature map a3 is the mask of the filter, and feature maps a1 and a2 are high-resolution feature maps after frequency fusion; The feature map output by the second feature extraction module of the backbone network and feature map a4 are input into the second frequency-domain perception fusion module of the frequency-domain perception feature fusion module, and feature maps b1, b2, and b3 are output. Feature map b1 and feature map b2 are added to obtain b4. Among them, feature map b3 is the mask of the filter, and feature maps b1 and b2 are high-resolution feature maps after frequency fusion; The feature map output by the first feature extraction module of the backbone network and feature map b4 are input into the third frequency-domain perception fusion module of the frequency-domain perception feature fusion module, and feature maps c1, c2, and c3 are output. Feature map c1 and feature map c2 are added to obtain feature map c4. Among them, feature map c3 is the mask of the filter, and feature maps c1 and c2 are high-resolution feature maps after frequency fusion; Feature map c4 is input into the first feature extraction module of the frequency-domain perception feature fusion module, and feature map F1 is output; Feature map F1 and feature map b4 are concatenated in the channel dimension and input into the second feature extraction module of the frequency-domain perception feature fusion module, and feature map F2 is output; Feature map F2 and feature map a3 are concatenated in the channel dimension and input into the third feature extraction module of the frequency-domain perception feature fusion module, and feature map F3 is output; Input the feature maps F1, F2, and F3 into the detection head to obtain the detection results.

[0037] Embodiment 2 As Figure 2 , Figure 5 shown, a method for detecting whether a judge wears a robe in a formal courtroom includes: S1: Collect image samples of a judge wearing a robe and not wearing a robe in a formal courtroom, cut the pictures with the minimum circumscribed rectangle of the judge's body contour and save them. Divide the cut pictures into two categories: wearing a robe and not wearing a robe, and construct a classification data set.

[0038] S2: Use the classification data set to train the resnet50 model.

[0039] S3: Collect real-time images of the formal courtroom scene, detect the judge's position through a person detection model, and extract the judge's coordinates in the image.

[0040] S4: Cut the regions where all detected judges exist with the minimum circumscribed rectangle of the judge's body contour, input the cut pictures into the trained Resnet50 model for classification, and determine whether all judges wear robes.

[0041] Take the classification result as the result of whether it is worn properly and then save it, that is, save each person's ID number, corresponding coordinate information, and whether they wear a robe.

[0042] S5: According to the extracted judge coordinates, track the position of the judge wearing a robe in real time, and judge whether the judge wears the robe properly by statistically analyzing the histogram information of the HSV three channels in the area where the judge is located.

[0043] As a refinement of the above embodiment, according to the extracted judge coordinates, track the position of the judge wearing a robe in real time, and judge whether the judge wears the robe properly by statistically analyzing the histogram information of the HSV three channels in the area where the judge is located. HSV refers to Hue, Saturation, and Value; the specific method is as follows: Input the extracted judge coordinates into the Deepsort algorithm for tracking, and assign numbers to the detected judges; Deepsort is an object tracking algorithm based on deep learning. Object tracking is to continuously locate and estimate the state of an object of interest in a video sequence to determine its position information in each frame, etc.; For the judge wearing a robe, according to the position tracked by the Deepsort algorithm, collect the position area every seconds during the court trial in real time, and statistically analyze the histogram information of the HSV three channels in this area; Calculate the Bhattacharyya distance between the histogram of the current moment and the corresponding channel histogram several seconds ago. After obtaining the Bhattacharyya distances of the three channels, calculate their mean value. If the mean value is less than the set threshold, it is determined that the judge is wearing a robe; if it is greater than the preset threshold, it indicates that the judge is not wearing a robe.

[0044] As a refinement of the above embodiment, the method for calculating the Bhattacharyya distance between the histogram of the current moment and the corresponding channel histogram several seconds ago is as follows: Obtain the histogram statistical information of the HSV three channels of the cut image area with the same number at the current moment and several seconds ago; Calculate the Bhattacharyya distance between the histograms of the corresponding channels at two moments. First, calculate the Bhattacharyya coefficient. Taking the H channel as an example, the formula is as follows: , where, represents the pixel-level range, represents the histogram of the H channel several seconds ago, represents the histogram of the H channel at the current moment, represents the Bhattacharyya coefficient; , where, represents the Bhattacharyya distance between the histogram of the H, S, or V channel several seconds ago and the histogram of the corresponding channel at the current moment.

[0045] The value range is any integer second between [1, 10]. The smaller m is, the more accurate the system detection result is. When the system has a high requirement for the detection accuracy of the judge's robe-wearing state, m should take a smaller value [1, 5] to increase the data acquisition frequency and thus more timely reflect the change of the judge's robe-wearing state. In scenarios where the system computing resources are limited or the real-time requirement is low, m can take a larger value [6, 10] to optimize the system performance and reduce resource consumption.

[0046] The method for selecting the set threshold is as follows: Collect the image samples of all judges wearing and not wearing robes within a continuous period of time. Take the image samples of the same-numbered judge wearing and not wearing a robe as a pair of samples, and extract the histogram information of the HSV three channels of the judge's area respectively; If there is only one pair of samples, calculate the Bhattacharyya distance between the histograms of the HSV three channels of the samples wearing and not wearing a robe; and multiply it by an adjustment coefficient as the set threshold; the value range of the adjustment coefficient α is (0, 1), and in this embodiment, it takes 0.8.

[0047] If there are multiple groups of sample pairs, calculate the Bhattacharyya distance for each group of samples to obtain a set of Bhattacharyya distance values; calculate the average of the Bhattacharyya distances between the HSV three-channel histograms of each sample pair and multiply it by an adjustment coefficient as the set threshold; the value range of the adjustment coefficient α is (0, 1), and 0.8 is taken in this embodiment. 20 to 30 judge image samples are taken in this embodiment.

[0048] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting the position of a judge in a courtroom, characterized in that, Including: Collect image samples of the courtroom with and without judges, construct a personnel detection dataset, and label the positions of the judges. Construct a personnel detection model based on yolov8s; the personnel detection model includes a backbone network, a frequency-domain perception feature fusion module, and a detection head. The outputs of the second convolution module, the third convolution module, and the fourth convolution module of the backbone network are all input into the feature extraction module after passing through the wavelet convolution module; the frequency-domain perception feature fusion module sequentially includes a first frequency-domain perception fusion module, a second frequency-domain perception fusion module, a third frequency-domain perception fusion module, a first feature extraction module, a second feature extraction module, and a third feature extraction module. Use the dataset to train the personnel detection model based on yolov8s. Collect the image of the judge's bench position in the courtroom and input it into the trained personnel detection model to obtain the position of the judge.

2. The method for detecting the position of a judge in a standard courtroom according to claim 1, wherein The backbone network of the personnel detection model based on yolov8s sequentially includes a first convolution module, a second convolution module, a first wavelet convolution, a first feature extraction module, a third convolution module, a second wavelet convolution, a second feature extraction module, a fourth convolution module, a third wavelet convolution, a third feature extraction module, a fifth convolution module, a fourth feature extraction module, and an SPPF module.

3. The method for detecting the position of a judge in a standard courtroom according to claim 2, wherein, The processing method of the feature map of the personnel detection model is as follows: Input the image into the backbone network; the feature map output by the SPPF module of the backbone network and the feature map output by the fourth feature extraction module are input into the first frequency-domain perception fusion module of the frequency-domain perception feature fusion module, and the output obtains feature maps a1, a2, and a3. Add feature map a1 and feature map a2 to obtain feature map a4. Among them, feature map a3 is the mask of the filter, and feature map a1 and feature map a2 are high-resolution feature maps after frequency fusion. The feature map output by the second feature extraction module of the backbone network and feature map a4 are input into the second frequency-domain perception fusion module of the frequency-domain perception feature fusion module, and the output obtains feature maps b1, b2, and b3. Add feature map b1 and feature map b2 to obtain feature map b4. Among them, feature map b3 is the mask of the filter, and feature map b1 and feature map b2 are high-resolution feature maps after frequency fusion. The feature map output by the first feature extraction module of the backbone network and feature map b4 are input into the third frequency-domain perception fusion module of the frequency-domain perception feature fusion module, and the output obtains feature maps c1, c2, and c3. Add feature map c1 and feature map c2 to obtain feature map c4. Among them, feature map c3 is the mask of the filter, and feature map c1 and feature map c2 are high-resolution feature maps after frequency fusion. Input feature map c4 into the first feature extraction module of the frequency-domain perception feature fusion module, and the output obtains feature map F1. Concatenate feature map F1 and feature map b4 in the channel dimension and input them into the second feature extraction module of the frequency-domain perception feature fusion module, and the output obtains feature map F2. Concatenate feature map F2 and feature map a3 in the channel dimension and input them into the third feature extraction module of the frequency-domain perception feature fusion module, and the output obtains feature map F3. Input the feature maps F1, F2, and F3 into the detection head to obtain detection results.

4. A detection method for regulating whether a judge wears a robe in court, characterized in that, Including: Collect image samples of a standard courtroom with judges wearing and not wearing robes. Crop and save the pictures using the minimum bounding rectangle of the judge's body contour. Divide the cropped pictures into two categories: wearing a robe and not wearing a robe, and construct a classification dataset. Use the classification dataset to train the resnet50 model. Collect real-time images of a standard courtroom scene. Detect the judge's position through the person detection model of the judge position detection method in any one of claims 1-3, and extract the judge's coordinates in the image. Crop all the detected regions where there are judges using the minimum bounding rectangle of the judge's body contour. Input the cropped pictures into the trained Resnet50 model for classification to determine whether all judges are wearing robes. Track the position of the judge determined to be wearing a robe in real time based on the extracted judge coordinates, and judge whether the judge is wearing the robe standardly by statistically analyzing the histogram information of the HSV three channels in the area where the judge is located.

5. The method for detecting whether a judge wears a robe in a formal court according to claim 4, characterized in that, Track the position of the judge determined to be wearing a robe in real time based on the extracted judge coordinates, and judge whether the judge is wearing the robe standardly by statistically analyzing the histogram information of the HSV three channels in the area where the judge is located. The specific method is as follows: Input the extracted judge coordinates into the Deepsort algorithm for tracking, and assign numbers to the detected judges. For the positions tracked by the judge in a robe according to the DeepSort algorithm, the position areas are collected in real time every seconds during the court trial, and the histogram information of the three HSV channels in this area is statistically analyzed; Calculate the Bhattacharyya distance between the histogram of the current moment and the corresponding channel histogram seconds ago. After obtaining the Bhattacharyya distances of the three channels, calculate their mean. If the mean is less than the set threshold, it is determined that the judge is wearing a robe; if it is greater than the preset threshold, it means that the judge is not wearing a robe.

6. The method for detecting whether a judge wears a robe in a standard courtroom according to claim 5, characterized in that Calculate the Bhattacharyya distance between the histogram of the current moment and the histogram of the corresponding channel seconds ago, and the specific method is as follows: Obtain the histogram statistical information of the HSV three channels of the sliced image area with the same number before seconds at the current moment; Calculate the Bhattacharyya distance between the histograms of corresponding channels at two moments. First, calculate the Bhattacharyya coefficient. The formula is as follows: , Among them, represents the pixel-level range, represents the histogram of the H, S, or V channel before seconds, represents the histogram of the H, S, or V channel at the current moment, represents the Bhattacharyya coefficient; Calculate the Bhattacharyya distance of all channels according to the Bhattacharyya distance formula. The Bhattacharyya distance formula is as follows: , Among them, represents the Bhattacharyya distance between the histogram of the H, S, or V channel before seconds and the histogram of this channel at the current moment.

7. The method for detecting whether a judge wears a robe in a standard courtroom according to claim 5, characterized in that, The said is any integer second within the range of [1, 10].

8. The method for detecting whether a judge wears a robe in a formal court according to claim 5, characterized in that The method for selecting the set threshold is as follows: Collect image samples of all judges wearing and not wearing robes within a continuous period of time. Take the image samples of the same numbered judge wearing and not wearing a robe as a pair of samples, and respectively extract the histogram information of the HSV three channels in the area where the judge is located. Calculate the Bhattacharyya distance between the HSV three-channel histograms of each pair of samples, take the average value of the Bhattacharyya distances of all pairs of samples, and multiply it by an adjustment coefficient as the set threshold. The value range of the adjustment coefficient α is (0, 1).

9. A detection device for regulating whether a judge wears a robe in a court, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the method for detecting whether a judge in a standard courtroom wears a robe as described in any one of claims 4-8 when running the program instructions.

10. A computer-readable storage medium, characterized in that, It stores a computer program, which when executed by the processor implements the method for detecting whether a judge in a standard courtroom wears a robe as described in any one of the above claims 4-8.

Citation Information

Patent Citations

  • Court trial inspection method and system

    CN110647831A

  • Method and system for detecting grape clusters in hard core period based on improved YOLOv5s

    CN117392529A

  • Factory safety wearing target detection method based on YOLOv7 improvement

    CN117496129A

  • Safety helmet standard wearing detection method based on improved YOLOv8

    CN119580177A

  • Unmanned aerial vehicle small target detection method based on YOLOv8 network

    CN119600261A

Cited By

  • Method for detecting whether judicial robe wearing of judge on court is standard or not

    CN120748013A

  • Classification model and feature fusion-based method for detecting behavior of replacing forensic gown in court

    CN122435692A