Electronic cemetery and informatization cemetery management method and system
By introducing local texture descriptors and motion consistency constraints in the background modeling process, the background misjudgment problem caused by vegetation movement in the cemetery virtual scene modeling is solved, and the authenticity of the virtual scene and the accuracy of static area segmentation are improved.
Patent Information
- Application Number
- CN202510249294.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
In the construction of smart cemetery, the virtual scene modeling of the cemetery is affected by non-rigid body movements such as vegetation, resulting in a high background misjudgment rate, affecting the authenticity of the scene construction.
In the background modeling process, local texture descriptors and motion consistency constraints are introduced. By comparing the similarity of local texture descriptors and calculating the differences in local motion consistency constraints, they distinguish vegetation areas and prospective goals, and reduce misjudgment.
It effectively suppresses background misjudgment caused by vegetation movement, improves the accuracy of static area segmentation, and enhances the authenticity of virtual scenes.
Smart Images

Figure CN120182909A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of virtual reality technology, and particularly to an electronic cemetery and an information-based cemetery management method and system. Background Art
[0002] In the construction of smart cemeteries, using computer vision technology to model and analyze the cemetery environment is an important direction. By processing and analyzing video surveillance data, the static background structure and dynamic personnel activity information of the cemetery environment can be obtained, a virtual scene can be constructed, and visual cemetery management can be realized. However, in practical applications, affected by the complexity of the cemetery environment, environmental modeling and virtual scene construction still face many challenges.
[0003] Among them, the existence of a large number of non-rigid objects such as trees and lawns in the cemetery poses difficulties for environmental modeling. These plants will have non-rigid movements under natural conditions such as wind and sunlight, such as the fluttering of leaves and the swaying of lawns. Traditional background modeling methods such as Gaussian mixture models, which are mainly based on the statistical distribution of pixel values, are difficult to accurately depict such complex background changes. The movement of plants will cause significant changes in pixel values, resulting in their misjudgment as foregrounds, thus affecting the accuracy of the background model. This problem of background misjudgment caused by plant movement will further affect the segmentation of static and dynamic areas in the cemetery and reduce the accuracy of environmental perception and state analysis.
[0004] For example, the related patent document CN117974394B discloses a virtual reality-based intelligent service system for digital tomb-sweeping in cemeteries. This invention relates to the field of virtual reality technology and solves the problem that the authenticity of the virtual scene will be affected if the dynamic monomers in the virtual scene are not confirmed. By determining the surrounding environments of different tombs and generating virtual scenes belonging to different tombs according to the specific pictures of different environments, static monomers or dynamic monomers are selected from the virtual scenes. For the selected dynamic monomers, their corresponding dynamic data are determined, and then based on the determined dynamic data, the dynamic monomers are made to move dynamically to ensure that the entire virtual scene can be in a dynamically running state, thereby ensuring the authenticity of the virtual scene and generating a real scene in combination with the actual monitoring pictures. However, this solution does not fully consider the problem of non-rigid movement of a large number of plants in the cemetery environment. Trees, lawns, etc. will have complex non-rigid movements under the action of wind, which can easily cause misjudgment in background modeling and affect the authenticity of scene construction. Summary of the Invention
[0005] In view of the high background misjudgment rate in the modeling of the virtual cemetery scene in the prior art due to the influence of non-rigid body movements such as vegetation, the present application provides an electronic cemetery and an information-based cemetery management method and system. By introducing local texture descriptors and motion consistency constraints in the background modeling process, the background misjudgment caused by vegetation movement is effectively suppressed, and the accuracy of static area segmentation is improved.
[0006] The purpose of the present application is achieved through the following technical solutions.
[0007] One aspect of the present application provides an electronic cemetery and an information-based cemetery management method, including: S1, collecting video surveillance data of the cemetery environment; S2, performing clustering analysis on the video surveillance data to obtain static areas and dynamic areas in the cemetery environment; S3, using an object detection algorithm to detect image frames in the dynamic area to obtain dynamic monomers, where the dynamic monomers include tourists and staff; S4, performing tracking detection on the dynamic monomers to obtain the motion trajectories of the dynamic monomers; S5, constructing a three-dimensional model of the static area according to the video surveillance data of the static area using three-dimensional reconstruction technology; S6, constructing a virtual scene of the cemetery environment according to the three-dimensional model of the static area and the motion trajectories of the dynamic monomers, and performing cemetery management according to the virtual scene.
[0008] Furthermore, the static area corresponds to fixed buildings in the cemetery; the dynamic area corresponds to moving targets in the cemetery.
[0009] Furthermore, S2, performing clustering analysis on the video surveillance data to obtain static areas and dynamic areas in the cemetery environment, includes: S21, performing foreground extraction on the video surveillance data to obtain a foreground target mask; the foreground target mask represents the pixel area with movement or change in the video frame; S22, using binary morphological operations to perform closing operation connection processing on the foreground target mask to obtain a moving target contour, and marking the moving target contour as the dynamic area; S23, using the complement of the foreground target mask as a non-moving area mask, and performing clustering analysis on the pixel points in the non-moving area mask using the K-means clustering algorithm to obtain K clustering clusters, where each clustering cluster corresponds to an initial semantic area; S24, using the region growing algorithm to optimize the initial semantic areas, and merging the initial semantic areas belonging to the same semantics into one area according to the spatial adjacency relationship and pixel value similarity between regions to obtain the static area.
[0010] Further, in S21, foreground extraction is performed on the video surveillance data to obtain a foreground object mask, including: for each pixel point in the video surveillance data, K Gaussian distributions are established; using the local binary pattern (LBP) operator or the histogram of oriented gradients (HOG) operator, the texture features within the neighborhood of each pixel point are extracted to obtain the local texture descriptor of the corresponding pixel point. Although there is non-rigid motion in the vegetation area, its local texture features usually remain stable. In this application, LBP or HOG is used to extract the texture features within the neighborhood of the pixel point to form a local texture descriptor. For the vegetation area, even if there is motion, the local texture descriptor thereof still has a high similarity with the local texture descriptors of the surrounding background pixels. Therefore, by comparing the similarity of the local texture descriptors, the vegetation area can be distinguished from the true foreground object, avoiding misjudging it as a dynamic area and improving the accuracy of background modeling.
[0011] The motion vector within the neighborhood of each pixel point is calculated using the optical flow method or the frame difference method to obtain the local motion consistency constraint of the corresponding pixel point. Although there is motion among the internal pixels of the vegetation area, it usually exhibits a cooperative and continuous motion pattern, such as the jitter of leaves and the undulation of the lawn. There are significant differences in motion consistency between this local cooperative motion and the independent motion of the true foreground object. In this application, the optical flow method or the frame difference method is used to calculate the motion vector within the neighborhood of the pixel to obtain the local motion consistency constraint. For the vegetation area, the motion vectors of its internal pixels have a high consistency in terms of direction and magnitude, while the true foreground object exhibits an independent motion pattern. By measuring the difference in the local motion consistency constraint, the vegetation area and the foreground object can be further distinguished, reducing misjudgment.
[0012] The EM algorithm is used to perform parameter estimation and update on the K Gaussian distributions of each pixel point to obtain a background model, and the background model reflects the statistical characteristics of the pixel points in the static area of the video frame; the Mahalanobis distance between the feature of each pixel point in the current video frame and the corresponding Gaussian distribution in the background model is calculated. If the Mahalanobis distance is greater than the distance threshold, the corresponding pixel point is marked as a preliminary foreground, otherwise it is marked as a background; for the pixel points marked as preliminary foreground, the similarity between the local texture descriptor of the pixel point and the local texture descriptors of the background pixel points in the corresponding neighborhood is calculated, and the difference between the local motion consistency constraint of the pixel point and the local motion consistency constraints of the background pixel points in the corresponding neighborhood is calculated; if the similarity between the local texture descriptors is greater than the texture threshold and the difference between the local motion consistency constraints is less than the motion threshold, the corresponding preliminary foreground pixel point is marked as a background; according to the foreground or background markings of all pixel points, a foreground object mask is obtained.
[0013] Among them, the Local Binary Patterns (LBP) operator is an operator for texture feature extraction. In foreground extraction, the LBP operator can be used to extract the texture features within the neighborhood of a pixel point, obtaining the local texture descriptor of the pixel point for distinguishing foreground and background regions.
[0014] The Histogram of Oriented Gradients (HOG) operator is an operator for extracting local gradient direction features of an image. In this application, the HOG operator can be used to extract the gradient direction features within the neighborhood of a pixel point, obtaining the local texture descriptor of the pixel point for distinguishing foreground and background regions.
[0015] The Expectation-Maximization (EM) algorithm is an algorithm for estimating model parameters through an iterative approach and is commonly used for parameter estimation of probability models containing latent variables. In this application, the EM algorithm is used to estimate the Gaussian distribution parameters of each pixel point. In the E-step, according to the current model parameters, the posterior probability (responsiveness) of each pixel point belonging to each Gaussian distribution is calculated as the latent variable. In the M-step, based on the features of the pixel points and the latent variable, the mean, covariance matrix, and weight parameters of the Gaussian distribution are updated by maximizing the likelihood function. By iteratively executing the EM algorithm, the background model of the pixel points, i.e., the parameter estimation result of the Gaussian mixture model, can be obtained. The EM algorithm can effectively handle the multimodal background of pixel points and adapt to the changes in dynamic backgrounds.
[0016] The foreground object mask is a binary image used to represent the pixel points belonging to the foreground object in a video frame. In the foreground object mask, the pixel with a value of 1 (white) represents the foreground object, and the pixel with a value of 0 (black) represents the background region. In this application, through background modeling and foreground / background decision-making for each pixel point, the foreground or background label of the pixel point is obtained to form the foreground object mask. This mask represents the position and range of the pixel points belonging to the moving object or changing region in the video frame and can be used for further processing and analysis.
[0017] Furthermore, the EM algorithm is used to estimate and update the parameters of the K Gaussian distributions for each pixel point to obtain the background model, including: initializing the parameters of the K Gaussian distributions for each pixel point in the current frame according to the parameters of the K Gaussian distributions of the pixel points in the previous video frame, where the parameters include the mean, variance, and weight; conventional EM algorithms usually adopt random initialization or experience-based initialization methods, while in this application, by using the modeling results of the previous frame to initialize the Gaussian distribution parameters of the current frame, the temporal continuity of the video sequence is fully utilized, obtaining a more reasonable and accurate initial estimate and accelerating the convergence speed of the algorithm. At the same time, by fusing local texture descriptors and motion consistency constraints, spatial and temporal context information is introduced, enhancing the adaptability of the background model to complex scenes and improving the accuracy and robustness of background modeling.
[0018] Calculate the likelihood probability of the features of each pixel point in the current video frame and the corresponding K Gaussian distributions to obtain the posterior probability of the pixel point belonging to each Gaussian distribution as the latent variable; the features include local texture descriptors and local motion consistency constraints; conventional EM algorithms mainly model based on the color information of pixel points, while in this application, when calculating the posterior probability of pixel points, color, texture, and motion information are fused, which helps to more accurately distinguish foreground and background regions. At the same time, by introducing local texture descriptors and motion consistency constraints, the background misjudgment caused by vegetation movement is effectively suppressed, improving the accuracy and reliability of foreground extraction.
[0019] Use the features and latent variables of each pixel point in the current video frame to update the parameters of the corresponding K Gaussian distributions through the maximum expectation algorithm; repeat the above steps of updating the Gaussian distribution parameters until the maximum number of iterations is reached to obtain the background model of the current video frame.
[0020] According to the background model, update the mean of the corresponding Gaussian distribution using the local texture descriptor of the background pixel point, and update the variance of the corresponding Gaussian distribution using the local motion consistency constraint of the background pixel point to obtain the updated background model. Conventional EM algorithms update the Gaussian distribution parameters by maximizing the likelihood function, mainly relying on the color information of pixel points. In this application, by adopting an update strategy based on the local texture descriptor and motion consistency constraint of the background pixel point, the background model can timely adapt to the dynamic changes of the scene, such as illumination changes and background motion, etc.
[0021] Further, in S3, use the object detection algorithm to detect the image frames in the dynamic area to obtain dynamic individuals, where the dynamic individuals include tourists and staff, including: S31, use the pre-trained convolutional neural network to extract features from the image frames in the dynamic area to obtain the feature map of the image frames; S32, use the Region Proposal Network (RPN) to generate candidate regions of the feature map; S33, use the Region of Interest (ROI) pooling layer to map candidate regions of different sizes into candidate region feature maps of a fixed size; S34, use the fully connected layer to classify and regress the candidate region feature maps, calculate the concept that the candidate regions belong to human bodies or backgrounds, and generate the final detection boxes; S35, perform non-maximum suppression on the generated detection boxes to obtain the human body detection results within the image frames as the dynamic individuals.
[0022] Further, in S4, perform tracking detection on the dynamic individuals to obtain the motion trajectories of the dynamic individuals, including: S41, establish the motion model and observation model of the tracking target, where the motion model describes the dynamic evolution process of the target state, and the observation model describes the relationship between the target state and the observed quantity; where the target state includes the position and speed of the target, and the observed quantity is the target detection result within consecutive image frames; S42, according to the estimated value of the target state in the previous frame and the motion model, predict the prior estimated value of the target state in the current frame; according to the target detection result in the current frame and the observation model, calculate the likelihood estimated value of the target state in the current frame; according to the estimated value of the state in the previous frame and the observed data in the current frame, through the calculation of the prior estimation and the likelihood estimation, obtain the posterior state estimated value in the current frame as the tracking output.
[0023] Specifically, the estimated value of the target state in the previous frame is jointly estimated through the motion model and the observation model. Specifically, the estimated value of the target state in the previous frame is the posterior estimated value obtained through Bayesian inference during the tracking process of the previous frame. In the Bayesian inference framework, the estimation process of the target state is usually divided into two steps: prediction step: according to the posterior estimated value of the target state in the previous frame and the motion model, predict the prior estimated value of the target state in the current frame. The motion model describes the dynamic evolution process of the target state, and by predicting the estimated value of the state in the previous frame, the prior estimated value of the state in the current frame is obtained. Update step: according to the observed quantity in the current frame (i.e., the target detection result) and the observation model, calculate the likelihood estimated value of the target state in the current frame. The observation model describes the relationship between the target state and the observed quantity, and by comparing the observed quantity in the current frame with the prior estimated value of the state, the likelihood estimated value of the state is obtained. Then, according to the prior estimated value of the state and the likelihood estimated value of the state, through the Bayesian inference formula, update the posterior estimated value of the target state in the current frame.
[0024] Therefore, the estimated value of the target state in the previous frame is actually jointly estimated through the motion model and the observation model, and it is the final output result in the tracking process of the previous frame. It synthesizes the prediction of the target state by the motion model and the correction of the target state by the observation model, representing the optimal estimate of the target state in the previous frame. In the tracking process of the current frame, this estimated value of the target state in the previous frame (i.e., the posterior estimated value of the previous frame) is used as the input for the prediction step of the current frame, and the state is predicted through the motion model to obtain the prior estimated value of the target state in the current frame. Then, the state is updated through the observed quantity and the observation model of the current frame to obtain the posterior estimated value of the target state in the current frame. This process makes full use of the correlation of time-series data and improves the accuracy and robustness of state estimation.
[0025] S43. Update the posterior estimated value of the target state in the current frame through Bayesian inference according to the prior estimated value of the target state and the likelihood estimated value of the target state, and use it as the tracking output.
[0026] S44. Calculate the similarity between the estimated values of the target states of the front and rear frames. If the similarity is greater than the threshold, connect the two estimated values of the target states into the same trajectory to obtain the motion trajectory of the dynamic monomer. By comparing the similarity between the estimated values of the states of the front and rear frames, the estimated values with high similarity are connected into the same trajectory to obtain the complete motion trajectory of the personnel in the cemetery. This method can automatically process a large amount of video surveillance data and quickly obtain the motion information of the personnel. Specifically, the estimated value of the target state is the posterior estimated value of the target state output in S43 because the posterior estimation is the optimal estimated value based on all observable information.
[0027] Preferably, a motion model and an observation model of the tracking target are established. The motion model describes the dynamic evolution process of the target state, and the observation model describes the relationship between the target state and the observed quantity. Among them, the target state includes the position, velocity, and scale of the target, and the observed quantity is the target detection result and the target appearance feature in consecutive image frames. According to the estimated value of the target state in the previous frame and the motion model, the prior estimated value of the target state in the current frame is predicted. According to the target detection result, the target appearance feature, and the observation model in the current frame, the likelihood estimated value of the target state in the current frame is calculated. According to the prior estimated value and the likelihood estimated value of the target state, the posterior estimated value of the target state in the current frame is updated through Bayesian inference and used as the tracking output. During the Bayesian inference process, an adaptive weight adjustment mechanism is introduced to dynamically adjust their weights in the posterior estimation according to the reliability of the prior estimated value and the likelihood estimated value of the target state to cope with the sudden change of target motion or the instability of the detection result. Calculate the similarity between the estimated values of the target state in the previous and current frames, considering the changes in the target position, velocity, and appearance features at the same time. If the similarity is greater than the first threshold, the two estimated values of the target state are connected as the same trajectory. If the similarity is less than the first threshold but greater than the second threshold, the current estimated value of the target state is marked as a possible trajectory break point and the trajectory is repaired in subsequent frames. If the similarity is less than the second threshold, the current estimated value of the target state is regarded as a new trajectory starting point. For the estimated value of the target state marked as a possible trajectory break point, the trajectory is repaired in multiple subsequent frames. By considering the motion consistency and appearance similarity of the target, search for possible trajectory connection paths in space and time and select the optimal connection path to repair the broken trajectory. If a suitable connection path cannot be found within the set time range, the trajectory at the break point is terminated and the subsequent estimated values of the target state are regarded as new trajectory starting points. In the Bayesian inference process of this application, an adaptive weight adjustment mechanism is introduced, which can dynamically adjust their weights in the posterior estimation according to the reliability of the prior estimated value and the likelihood estimated value of the target state, so as to better cope with the sudden change of target motion or the instability of the detection result. During the trajectory generation process, by introducing multiple similarity thresholds and a trajectory repair mechanism, the problem of trajectory breakage can be handled more flexibly. For possible trajectory break points, by repairing the trajectory in subsequent frames and considering the motion consistency and appearance similarity of the target, the optimal trajectory connection path can be found to repair the broken trajectory. By setting the time range for trajectory repair, the influence of long-term trajectory breakage on the continuity of tracking can be avoided. If a suitable connection path cannot be found within the set time range, the trajectory at the break point is terminated and the subsequent estimated values of the target state are regarded as new trajectory starting points, ensuring the continuity and effectiveness of the tracking process.
[0028] Further, in S6, according to the three-dimensional model of the static area and the movement trajectory of the dynamic monomer, a virtual scene of the cemetery environment is constructed, including: S61, loading the three-dimensional model of the static area using a three-dimensional rendering engine to construct the static background of the virtual scene; S62, matching the detection result of the dynamic monomer using a three-dimensional human model library, instantiating the three-dimensional human model corresponding to the dynamic monomer, and rendering the human model into the virtual scene according to the position of the dynamic monomer; S63, generating the movement animation of the human model in the virtual scene according to the movement trajectory of the dynamic monomer; S64, using the scene graph management of the three-dimensional rendering engine to perform fusion rendering on the static background and the dynamic human model to generate the virtual scene of the cemetery environment.
[0029] Further, in S63, generating the movement animation of the human model in the virtual scene according to the movement trajectory of the dynamic monomer includes: converting the position points on the movement trajectory of the dynamic monomer into the three-dimensional space coordinates of the human model in the virtual scene coordinate system as the key frames of the human model; calculating the number of interpolation frames and the interpolation coefficient between adjacent key frames according to the position change between adjacent key frames of the human model; using the linear interpolation algorithm to generate the three-dimensional space coordinate sequence of the interpolation frames according to the three-dimensional space coordinates and the interpolation coefficient of adjacent key frames; constructing the movement frame sequence of the human model according to the three-dimensional space coordinate sequences of the key frames and the interpolation frames; loading the movement frame sequence of the human model in the three-dimensional rendering engine to control the movement of the human model in the virtual scene.
[0030] Another aspect of the present application also provides an electronic cemetery and an information-based cemetery management system for executing an electronic cemetery and an information-based cemetery management method of the present application.
[0031] Compared with the prior art, the advantages of the present application are as follows:
[0032] The traditional Gaussian mixture background model is only based on the statistical distribution of pixel values and is difficult to depict complex background changes. In the modeling process of the present application, local texture descriptors and motion consistency constraints are fused and used as the feature inputs of pixel points into the Gaussian mixture model to guide the estimation and update of model parameters. By combining local spatial information and motion information, the Gaussian distribution can better adapt to the non-rigid motion of vegetation while maintaining sensitivity to real foreground targets. This fusion strategy can establish a robust background model, accurately segmenting the dynamic and static areas while avoiding the interference of vegetation movement.
[0033] During the foreground extraction process, the Mahalanobis distance between a pixel and the background model is used to preliminarily determine whether it belongs to the foreground. For the pixels preliminarily determined to be foreground, the present application further compares the similarity between its local texture descriptor and the background pixels in the neighborhood, as well as the difference between the local motion consistency constraint and the background pixels in the neighborhood. If the texture similarity is high and the motion consistency difference is small, then the pixel is relabeled as background, regarded as a misjudgment caused by vegetation movement. Through this correction mechanism, the residual vegetation movement area can be removed, and the accuracy of the foreground target mask can be further improved.
[0034] To adapt to the dynamic changes of the background, the background model needs to be continuously updated. During the update process, the present application uses the local texture descriptor of the current background pixel to update the mean value of the Gaussian distribution, and uses its local motion consistency constraint to update the variance of the Gaussian distribution. By introducing local features into the update mechanism, the background model can better adapt to the non-rigid motion of vegetation while maintaining the discriminative ability for foreground targets. This adaptive update strategy can maintain the stability and robustness of the background model and reduce the long-term impact of vegetation movement on background modeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The present application will be further described in the form of exemplary embodiments, which will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:
[0036] Figure 1 is an exemplary flowchart of an electronic cemetery and information-based cemetery management method shown in some embodiments of the present application;
[0037] Figure 2 is an exemplary flowchart of generating a dynamic area and a static area shown in some embodiments of the present application;
[0038] Figure 3 is an exemplary flowchart of generating a dynamic monomer shown in some embodiments of the present application;
[0039] Figure 4 is an exemplary flowchart of generating the motion trajectory of a dynamic monomer shown in some embodiments of the present application;
[0040] Figure 5 is an exemplary flowchart of generating a virtual scene shown in some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The methods and systems provided in the embodiments of the present application will be described in detail below with reference to the drawings.
[0042] As Figure 1As shown in the figure, collect the video surveillance data of the cemetery environment; perform clustering analysis on the video surveillance data to obtain the static areas and dynamic areas in the cemetery environment; use the object detection algorithm to detect the image frames in the dynamic areas to obtain dynamic entities, where the dynamic entities include visitors and staff; perform tracking detection on the dynamic entities to obtain the movement trajectories of the dynamic entities; according to the video surveillance data of the static areas, use the 3D reconstruction method to construct the 3D model of the static areas; according to the 3D model of the static areas and the movement trajectories of the dynamic entities, construct the virtual scene of the cemetery environment, and perform cemetery management according to the virtual scene.
[0043] S1. Collect the video surveillance data of the cemetery environment; deploy multiple video surveillance cameras in the cemetery environment to comprehensively cover different areas of the cemetery. Each camera collects high-definition video data in real time and transmits the video data to the central server for storage and processing. The video data contains a continuous sequence of image frames, reflecting the dynamic change process of the cemetery environment.
[0044] As Figure 2 shown in the figure, S2. Perform clustering analysis on the video surveillance data to obtain the static areas and dynamic areas in the cemetery environment, including: S21. Extract the foreground of the video surveillance data to obtain the foreground object mask: For each frame image in the video sequence, for each pixel point therein, establish K Gaussian distributions (usually K takes 3 to 5) as the background model. Each Gaussian distribution is described by three parameters: mean μ, covariance matrix Σ, and weight ω, which respectively represent the average value, variance, and proportion of the pixel point under this Gaussian distribution. Initially, set the mean of all Gaussian distributions to the color value of the corresponding pixel point, the covariance matrix to the identity matrix, and the weight to 1 / K.
[0045] Extract the local texture descriptor: For each pixel point, take this pixel point as the center and select a neighborhood window of a fixed size (such as 3×3 or 5×5). In this window, use the LBP operator or HOG operator to extract the local texture features of the pixel point. LBP operator: Compare each pixel value in the neighborhood window with the center pixel value. If the neighborhood pixel value is greater than the center pixel value, set the corresponding binary bit to 1, otherwise set it to 0. Connect all binary bits in a clockwise or counterclockwise order to obtain the LBP code. The LBP code can describe the local texture pattern of the area where the pixel point is located. HOG operator: In the neighborhood window, calculate the gradient magnitude and direction of each pixel point. Quantize the gradient direction into a preset direction histogram, and each direction corresponds to a histogram bin. Accumulate the gradient magnitudes falling into each bin to obtain the local gradient direction histogram as the HOG feature descriptor of the pixel point. The extracted LBP code or HOG feature vector is used as the local texture descriptor T of the pixel point.
[0046] Calculate the local motion consistency constraint: For two adjacent frames of images I t and I t+1 in the video sequence, use the optical flow method or the frame difference method to calculate the motion vectors of pixel points. Optical flow method: By estimating the motion velocity vector (u, v) of pixel points in the spatio-temporal domain, a dense optical flow field is obtained. For pixel point p(x, y), its velocity vector is (u(x, y), v(x, y)), indicating the motion displacement of this pixel point from the previous frame to the current frame. Frame difference method: Subtract two adjacent frames of images to obtain a difference image ΔI = I t+1 -I t . The non-zero pixels in the difference image indicate the existence of motion or change. Binarize the difference image to obtain a motion region mask. For the non-zero pixels in the mask, the change amount of their position coordinates is the motion vector (Δx, Δy). The calculated motion vector (u, v) or (Δx, Δy) is used as the local motion consistency constraint M of the pixel point.
[0047] Estimate the Gaussian background model parameters using the EM algorithm: Use the EM algorithm to estimate and update the parameters of K Gaussian distributions for each pixel point to establish a background model. The EM algorithm alternately executes the E step and the M step to estimate the parameters of the Gaussian distribution. E step: According to the currently estimated Gaussian distribution parameters, calculate the posterior probability P(k|x) of each pixel point belonging to each Gaussian distribution as the hidden variable. Among them, x represents the feature vector of the pixel point, including the color value, the local texture descriptor T, and the local motion consistency constraint M. M step: Use the feature vector x of the pixel point and the hidden variable P(k|x) to update the parameters (μ, Σ, ω) of the Gaussian distribution through the maximum expectation algorithm. Iteratively execute the EM algorithm until convergence or the maximum number of iterations is reached to obtain the background model of the pixel point.
[0048] When processing the current frame of image, use the Gaussian distribution parameters (μ, Σ, ω) of the pixel points in the previous frame as the initial parameters of the corresponding pixel points in the current frame. For each pixel point in the current frame, extract its color feature x c , local texture descriptor T, and local motion consistency constraint M to form the feature vector x. Calculate the likelihood probability p(x|k) of x and each Gaussian distribution in the background model to obtain the posterior probability P(k|x) of the pixel point belonging to each Gaussian distribution.
[0049] Use the local texture descriptor T of the background pixel points to update the mean μ k of the corresponding Gaussian distribution, and use the local motion consistency constraint M of the background pixel points to update the covariance matrix Σ k of the corresponding Gaussian distribution to make the background model adapt to the dynamic changes of the scene. The update formula is: μ k ' = (1 - α)×μ k + α×T, Σ k ' = (1 - α)×Σk +α×(M - μ k )×(M - μ k ) T , where α is the learning rate that controls the speed of background model update. Here, μ k is the mean vector of the k-th Gaussian distribution, representing the central position of the k-th Gaussian distribution. It is a d-dimensional vector, where d is the dimension of the feature space. T is the local texture descriptor of the current background pixel point, representing the texture features within the neighborhood of the current pixel point. It is a d-dimensional vector with the same dimension as μ k . α is the learning rate that controls the speed of background model update. It is a scalar with a value range of (0, 1]. A larger α value indicates a faster update of the background model, while a smaller α value indicates a slower update of the background model. Σ k is the covariance matrix of the k-th Gaussian distribution, representing the shape and direction of the k-th Gaussian distribution. It is a d×d matrix, where d is the dimension of the feature space. M is the local motion consistency constraint of the current background pixel point, representing the motion features within the neighborhood of the current pixel point. It is a d-dimensional vector with the same dimension as μ k . (M - μ k ) is the difference vector between the local motion consistency constraint of the current background pixel point and the mean of the k-th Gaussian distribution. It represents the degree of deviation between the motion features of the current pixel point and the central position of the k-th Gaussian distribution. (M - μ k ) T The transpose of the difference vector (M - μ k ) is used to calculate the outer product of the covariance matrix.
[0050] For each pixel point in the current frame, calculate the Mahalanobis distance d M (x, k) between its feature vector x and the corresponding Gaussian distribution in the background model. The Mahalanobis distance measures the degree of difference between the pixel point features and the background model, and the calculation formula is: d M (x, k) = (x - μ k ) T × Σ k -1 × (x - μ k ). If the Mahalanobis distance d M (x, k) is greater than the set threshold T d , then mark the pixel point as a preliminary foreground, otherwise mark it as a background.
[0051] For the pixel points marked as preliminary foreground, extract their local texture descriptors T f and local motion consistency constraints M f . Calculate the similarity S b between T_f and the local texture descriptors T T of the background pixel points in the neighborhood, as well as Mf The local motion consistency constraint M with background pixel points in the neighborhood b The difference D M If the similarity S T is greater than the set texture threshold T T and the difference D M is less than the set motion threshold T M then the preliminary foreground pixel points are relabeled as background, regarded as misjudgments caused by vegetation movement. According to the foreground or background labels of all pixel points, a binary foreground object mask image is generated. The pixel points labeled as foreground are set to 1, and the pixel points labeled as background are set to 0 to obtain the foreground object mask. This application utilizes the color, texture, and motion information of pixel points and performs adaptive background modeling and updating through the EM algorithm, which can effectively cope with the illumination changes and dynamic background interference in the scene, and improve the robustness and accuracy of foreground object detection.
[0052] S22, Use binary morphological operations to perform closing operation connection processing on the foreground object mask: Perform morphological closing operation on the foreground object mask image. Through the operations of dilation first and then erosion, small holes in the foreground area are filled, and the fragmented foreground areas are connected to obtain a complete moving object contour. The extracted connected region contour is labeled as the dynamic region mask.
[0053] S23, Take the complement of the foreground object mask to obtain the non-moving region mask, and perform clustering analysis using the K-means clustering algorithm: For each pixel point in the foreground object mask image, if its value is 1 (foreground), it is set to 0; if its value is 0 (background), it is set to 1. The complement image of the foreground object mask is obtained, representing the non-moving region mask, corresponding to the static background region. For each pixel point in the non-moving region mask, its color features (such as RGB values, HSV values, etc.) and texture features (such as LBP features, Gabor features, etc.) are extracted. The extracted features are composed into a feature vector to represent the feature description of each pixel point. Specify the number of clustering clusters K. Usually, a suitable K value is selected according to the scene complexity and the number of required semantic region categories. Randomly select K pixel points as the initial clustering centers. Traverse all non-moving region pixel points, calculate the Euclidean distance between each pixel point and the K clustering centers, and assign the pixel point to the nearest clustering cluster. For each clustering cluster, calculate the mean value of the features of all pixel points within the cluster, and update the mean value as the new clustering center. Repeat the above two steps until the clustering centers no longer change significantly or reach the maximum number of iterations. The label assignment of the clustering cluster to which each pixel point belongs is assigned to the corresponding pixel point to obtain the initial semantic segmentation result. The K-means clustering algorithm divides the non-moving region pixel points into K clustering clusters through iterative optimization, and each clustering cluster corresponds to an initial semantic region, such as buildings, roads, greenery, etc.
[0054] S24. Optimize the initial semantic regions using the region growing algorithm to obtain the final static regions: Using the initial semantic regions obtained by K-means clustering as seeds, perform region merging using the region growing algorithm. The region growing algorithm takes into account the spatial adjacency relationship and pixel value similarity between regions, and gradually merges adjacent initial regions belonging to the same semantics to obtain complete semantic regions. The merged regions are the static regions in the cemetery environment, corresponding to fixed facilities such as buildings and roads. Morphological closing operation is used to connect fragmented moving target regions, K-means clustering is used to obtain initial semantic regions, and the region growing algorithm is used to optimize and merge semantic regions, finally obtaining the static regions in the cemetery environment.
[0055] As Figure 3 shown, S3. Detect the image frames in the dynamic region using the object detection algorithm to obtain dynamic entities, where the dynamic entities include tourists and staff, including: S31. Use a pre-trained convolutional neural network (such as VGG, ResNet, etc.) to extract features from the image frames in the dynamic region. Input the image frames into the convolutional neural network, and through the calculations of multiple convolutional layers and pooling layers, obtain the feature maps of the image frames. The feature maps usually have a smaller spatial size and more channels, representing the high-level semantic features of the image frames.
[0056] S32. Slide a small convolutional network on the feature map to generate candidate regions. For each position on the feature map, the convolutional network outputs k candidate regions (called anchor boxes) with different scales and aspect ratios. At the same time, the convolutional network outputs the probability score indicating whether each candidate region contains the target object and the coordinate correction values of the candidate regions. Based on the probability scores and coordinate correction values, filter out high-quality candidate regions as the input for subsequent processing.
[0057] S33. Map the candidate regions of different sizes generated by RPN to a feature map of a fixed size. For each candidate region, perform a pooling operation on the feature map according to its coordinates to aggregate the features within the candidate region into a feature vector of a fixed size. Commonly used pooling operations include max pooling and average pooling. The purpose of pooling is to make the candidate regions have the same feature dimension for subsequent classification and regression.
[0058] S34. Input the feature map of the fixed-size candidate regions obtained by ROIPooling into the fully connected layer. The fully connected layer performs classification and regression calculations on the feature map of the candidate regions. The classification part calculates the probability that the candidate region belongs to a human body or the background through the Softmax activation function to determine whether the candidate region contains the target object. The regression part calculates the coordinate correction values of the candidate regions through a linear function to fine-tune the position and size of the candidate regions and generate more accurate detection boxes.
[0059] S35. Post-process the detected bounding boxes obtained from classification and regression to remove redundant and overlapping detection results. Sort the detected bounding boxes in descending order according to their classification probability scores. Select the bounding box with the highest score as the reference, and calculate the intersection over union (IoU) between it and other bounding boxes. If the IoU between another bounding box and the reference box is greater than a set threshold (such as 0.5), remove it, considering it as a duplicate detection result. Repeat the above steps until all bounding boxes are processed to obtain the final human detection result, i.e., the dynamic monomer.
[0060] As Figure 4 shown, S4. Perform tracking detection on the dynamic monomer to obtain the motion trajectory of the dynamic monomer, including: S41. Establish a motion model and an observation model for the tracking target. The motion model describes the dynamic evolution process of the target state. Commonly used motion models include the constant velocity model, the constant acceleration model, etc. Taking the constant velocity model as an example, assume that the target moves at a constant speed between two frames, and the state vector includes the position and velocity of the target. State transition equation: X(t) = A × X(t - 1) + W(t), where X(t) is the state vector of the current frame, A is the state transition matrix, and W(t) is the process noise. The observation model describes the relationship between the target state and the observed quantity. Commonly used observation models include the linear Gaussian model. Observation equation: Z(t) = H × X(t) + V(t), where Z(t) is the observed quantity of the current frame (such as the target detection result), H is the observation matrix, and V(t) is the observation noise; t is the current time step.
[0061] S42. According to the estimated value X(t - 1) of the target state in the previous frame and the motion model, predict the prior estimated value X(t)' of the target state in the current frame. Prediction equation: X(t)' = A × X(t - 1), indicating that the state transition matrix A is used to predict the estimated value of the state in the previous frame. According to the target detection result Z(t) in the current frame and the observation model, calculate the likelihood estimated value P(Z(t)|X(t)) of the target state in the current frame. The likelihood estimated value represents the probability of observing the target detection result Z(t) under the condition of the given state X(t).
[0062] S43. According to the prior estimated value X(t)' of the target state and the likelihood estimated value P(Z(t)|X(t)) of the target state, update the posterior estimated value X(t) of the target state in the current frame through Bayesian inference. Bayesian inference formula:
[0063] where P(X(t)|Z(t)) is the posterior probability, P(X(t)) is the prior probability, and P(Z(t)) is the normalization factor. Through methods such as maximizing the posterior probability or expected value estimation, obtain the estimated value X(t) of the target state in the current frame as the tracking output.
[0064] In S44, the similarity between the estimated target states X(t - 1) and X(t) of the previous and current frames is calculated. Common similarity metrics include Euclidean distance, Mahalanobis distance, etc. If the similarity is greater than the set threshold, it is considered that the two state estimates before and after belong to the same target, and they are connected into the same trajectory. The above process is repeated to connect the estimated target states of each frame to generate the complete motion trajectory of the dynamic monomer. Post-processing such as smoothing and breakpoint connection is performed on the generated trajectory to obtain the final motion trajectory of the dynamic monomer. In this application, the target tracking algorithm is used to track and detect the dynamic monomer, and the motion trajectory of the dynamic monomer is obtained. The motion model and the observation model respectively describe the dynamic evolution process of the target state and the relationship between the observed quantity and the state. Bayesian inference updates the posterior estimate of the target state by fusing the prior estimate and the likelihood estimate, realizing the recursive estimation and tracking of the target state. The similarity calculation and trajectory generation steps connect the estimated target states of different frames into a complete motion trajectory, describing the motion process of the dynamic monomer in the spatio-temporal domain.
[0065] As Figure 5 shown, in S5, based on the video surveillance data of the static area, a three-dimensional model of the static area is constructed using three-dimensional reconstruction technology; specifically, the video surveillance data of the static area is preprocessed, such as image enhancement, denoising, etc., to improve the image quality. Key frames are extracted from consecutive image frames, and frames with representativeness and large amounts of information are selected as the input for three-dimensional reconstruction. Feature extraction is performed on the selected key frame images. Common features include SIFT, SURF, ORB, etc. By matching the feature descriptors, corresponding points between different image frames are found, and the corresponding relationship between the image frames is established. Using the epipolar geometry constraint or the homography matrix, the matching points are filtered and optimized to remove incorrect matches.
[0066] Based on the matching points, the internal parameters (such as focal length, principal point) and external parameters (such as rotation matrix, translation vector) of the camera are estimated. Common camera parameter estimation methods include direct linear transformation (DLT), random sample consensus (RANSAC), etc. By minimizing the reprojection error or other error metrics, the camera parameters are optimized to obtain more accurate estimated values. Using the estimated camera parameters, the coordinates of the matching points in the three-dimensional space are calculated based on the triangulation principle to obtain a sparse point cloud. The sparse point cloud is filtered and denoised to remove outliers and abnormal points to improve the quality of the point cloud.
[0067] Based on the sparse point cloud, use multi-view stereo vision (MVS) algorithms, such as patch-based MVS, depth map fusion, etc., to generate a dense point cloud. By optimizing and refining the point cloud, such as surface meshing, normal vector estimation, etc., a more complete and smooth point cloud model is obtained. Map the texture information of the key frame image onto the point cloud or mesh model to obtain a three-dimensional model with texture. Simplify and optimize the mesh model, such as mesh decimation, mesh smoothing, etc., to reduce the complexity of the model and improve the rendering efficiency.
[0068] S6. Based on the three-dimensional model of the static area and the motion trajectory of the dynamic entity, construct a virtual scene of the cemetery environment, including: S61. Use a three-dimensional rendering engine, such as Unity3D, Unreal Engine, etc., to create a virtual scene. Import the three-dimensional model file of the static area into the three-dimensional rendering engine, such as model files in obj, fbx, etc. formats. According to the coordinate system and scale of the three-dimensional model, place it at an appropriate position and orientation in the virtual scene. Set the materials and textures of the three-dimensional model to make it have a realistic visual effect. Set the lighting, shadow and other effects of the virtual scene to create a realistic environmental atmosphere.
[0069] S62. According to the detection results of the dynamic entity, such as bounding boxes, key points, etc., determine the position and pose of the human body in the image. Use a pre-established three-dimensional human body model library, such as MakeHuman, DAZ3D, etc., to select a human body model that matches the dynamic entity. According to the position and pose of the dynamic entity, perform transformations such as scaling, rotation and translation on the human body model to align it with the position of the human body in the image. Import the transformed human body model into the virtual scene and integrate it with the static background.
[0070] S63. Generate the motion animation of the human body model. The motion trajectory of the dynamic entity is usually represented in the image coordinate system or the world coordinate system, and it needs to be converted into a sequence of three-dimensional space coordinates in the virtual scene coordinate system. Establish the conversion relationship between the image coordinate system, the world coordinate system and the virtual scene coordinate system, usually obtained by methods such as calibration or corresponding point matching. For each position point on the motion trajectory, use the coordinate conversion relationship to convert it from the image coordinate system or the world coordinate system to the three-dimensional space coordinates in the virtual scene coordinate system.
[0071] In the converted three-dimensional space coordinate sequence, select key frames as the key positions and postures of human motion. The selection of key frames can be determined according to the characteristics and requirements of the motion trajectory. Common key frames include: starting frame: the starting position of the motion trajectory. ending frame: the ending position of the motion trajectory. turning point: the position where the speed or direction changes significantly in the motion trajectory. equally spaced frames: frames selected at fixed time intervals on the motion trajectory. The selection of key frames should represent the key states of human motion as accurately as possible, while avoiding selecting too many or too few key frames to balance the fluency of the animation and the data volume.
[0072] For two adjacent key frames, calculate the position change amount and time interval between them. The position change amount can be obtained by calculating the Euclidean distance or other distance metrics between the key frames, representing the displacement of the human body between adjacent key frames. The time interval can be calculated based on the timestamps or frame numbers of the key frames, representing the time difference between adjacent key frames.
[0073] Based on the position change amount and time interval between adjacent key frames, calculate the number of interpolation frames to be inserted and the interpolation coefficients. The number of interpolation frames can be determined according to the time interval and the preset frame rate. For example, if the time interval is 0.5 seconds and the frame rate is 30 frames per second, then 15 interpolation frames need to be inserted. The interpolation coefficient represents the position ratio of the interpolation frame between adjacent key frames and can be calculated by linear interpolation or other interpolation methods.
[0074] Using the linear interpolation algorithm, generate the three-dimensional space coordinates of the interpolation frames based on the three-dimensional space coordinates of adjacent key frames and the interpolation coefficients. For each interpolation frame, calculate its interpolation coefficient between adjacent key frames, and then calculate the three-dimensional space coordinates of the interpolation frame through the linear interpolation formula: Let the three-dimensional space coordinates of adjacent key frames be P1 and P2, and the interpolation coefficient be t (0 <= t <= 1), then the three-dimensional space coordinates P of the interpolation frame can be expressed as: P = P1 + t × (P2 - P1). Repeat the above steps to generate the three-dimensional space coordinates of all interpolation frames. Combine the three-dimensional space coordinates of key frames and interpolation frames in chronological order to obtain the complete motion frame sequence of the human model. Each frame in the motion frame sequence corresponds to a position and posture of the human model in the virtual scene, and the transition between frames represents the motion of the human model. The frame rate of the motion frame sequence should be consistent with the rendering frame rate of the virtual scene to ensure the fluency of the animation.
[0075] In a 3D rendering engine, load the motion frame sequence of the human model. According to each frame in the motion frame sequence, set the position and pose of the human model in the virtual scene. Utilize the animation control functions of the 3D rendering engine, such as keyframe animation, skeletal animation, etc., to control the human model to move in the virtual scene according to the motion frame sequence. Through techniques such as interpolation and smoothing, make the movement of the human model more natural and smooth, avoiding abrupt or jittery phenomena.
[0076] S64. Integrate and render the virtual scene: In a 3D rendering engine, create a scene graph to manage all objects in the virtual scene. Add the static background model and the dynamic human model to the scene graph, and set their parent-child relationships and relative positions. According to the viewing angle and camera parameters, set the position and direction of the virtual camera to control the rendering perspective of the scene. Utilize the rendering pipeline of the 3D rendering engine to render all objects in the scene graph to generate the final virtual scene image. Perform post-processing on the virtual scene, such as adding special effects, adjusting color and contrast, etc., to enhance the visual effect.
Claims
1. An electronic cemetery and information-based cemetery management method, characterized in that: include: S1, collects video surveillance data of the cemetery environment; S2, cluster analysis is performed on the video surveillance data to obtain static areas and dynamic areas in the cemetery environment; S3, using a target detection algorithm to detect image frames in the dynamic area to obtain dynamic monomers, where the dynamic monomers include tourists and staff; S4, tracking and detecting the dynamic monomer to obtain the motion trajectory of the dynamic monomer; S5, constructing a three-dimensional model of the static area using a three-dimensional reconstruction method according to the video surveillance data of the static area; S6, constructing a virtual scene of the cemetery environment according to the three-dimensional model of the static area and the motion trajectory of the dynamic monomer, and managing the cemetery according to the virtual scene.
2. The electronic cemetery and information-based cemetery management method according to claim 1 is characterized by: The static area corresponds to the fixed buildings in the cemetery; The dynamic area corresponds to a moving object in the cemetery.
3. The electronic cemetery and information-based cemetery management method according to claim 2 is characterized by: S2, cluster analysis is performed on the video surveillance data to obtain static areas and dynamic areas in the cemetery environment, including: S21, performing foreground extraction on the video surveillance data to obtain a foreground target mask; the foreground target mask represents a pixel area with motion or change in the video frame; S22, using binary morphological operations to perform closed operation and connection processing on the foreground target mask to obtain the contour of the moving target, and marking the contour of the moving target as a dynamic area; S23, taking the complement of the foreground target mask as the non-motion region mask, and performing cluster analysis on the pixels in the non-motion region mask using the K-means clustering algorithm to obtain K clusters, each of which corresponds to an initial semantic region; S24, optimizing the initial semantic regions by using a region growing algorithm, merging the initial semantic regions belonging to the same semantics into one region according to the spatial adjacency relationship and pixel value similarity between the regions, and obtaining a static region.
4. The electronic cemetery and information-based cemetery management method according to claim 3 is characterized by: S21, extracting the foreground of the video surveillance data to obtain a foreground target mask, including: For each pixel in the video surveillance data, K Gaussian distributions are established; The local binary model LBP operator or the directional gradient histogram HOG operator is used to extract the texture features in the neighborhood of each pixel and obtain the local texture descriptor of the corresponding pixel. The motion vector in the neighborhood of each pixel is calculated using the optical flow method or the frame difference method to obtain the local motion consistency constraint of the corresponding pixel; The EM algorithm is used to estimate and update the parameters of K Gaussian distributions of each pixel to obtain a background model, wherein the background model reflects the statistical characteristics of the pixels in the static area of the video frame; Calculate the Mahalanobis distance between the feature of each pixel in the current video frame and the corresponding Gaussian distribution in the background model. If the Mahalanobis distance is greater than the distance threshold, mark the corresponding pixel as a preliminary foreground, otherwise mark it as background. For the pixel marked as preliminary foreground, the similarity between the local texture descriptor of the pixel and the local texture descriptor of the background pixel in the corresponding neighborhood is calculated, and the difference between the local motion consistency constraint of the pixel and the local motion consistency constraint of the background pixel in the corresponding neighborhood is calculated; If the similarity between local texture descriptors is greater than the texture threshold, and the difference between local motion consistency constraints is less than the motion threshold, the corresponding preliminary foreground pixel is marked as background; According to the foreground or background label of all pixels, the foreground target mask is obtained.
5. The electronic cemetery and information-based cemetery management method according to claim 4 is characterized by: The EM algorithm is used to estimate and update the parameters of the K Gaussian distributions of each pixel to obtain the background model, including: Initialize the K Gaussian distribution parameters of each pixel in the current frame according to the K Gaussian distribution parameters of the pixel in the previous frame of video, wherein the parameters include mean, variance and weight; Calculate the likelihood probability of the features of each pixel in the current video frame and the corresponding K Gaussian distributions, and obtain the posterior probability that the pixel belongs to each Gaussian distribution as a hidden variable; the features include local texture descriptors and local motion consistency constraints; Using the features and latent variables of each pixel in the current video frame, the parameters of the corresponding K Gaussian distributions are updated through the maximum expectation algorithm; Repeat the above steps of updating Gaussian distribution parameters until the maximum number of iterations is reached to obtain the background model of the current video frame; According to the background model, the local texture descriptor of the background pixels is used to update the mean of the corresponding Gaussian distribution, and the local motion consistency constraint of the background pixels is used to update the variance of the corresponding Gaussian distribution to obtain the updated background model.
6. The electronic cemetery and information-based cemetery management method according to claim 4 or 5, characterized in that: S3, using a target detection algorithm to detect image frames in the dynamic area to obtain dynamic monomers, wherein the dynamic monomers include tourists and staff, including: S31, extracting features of the image frame in the dynamic region using a pre-trained convolutional neural network to obtain a feature map of the image frame; S32, using the region candidate network RPN to generate a candidate region of the feature map; S33, using a region of interest pooling layer ROI to map candidate regions of different sizes into a candidate region feature map of a fixed size; S34, using the fully connected layer to classify and regress the candidate region feature map, calculate the concept of the candidate region belonging to the human body or the background, and generate the final detection frame; S35, performing non-maximum suppression on the generated detection frame to obtain a human body detection result in the image frame as a dynamic monomer.
7. The electronic cemetery and information-based cemetery management method according to claim 6 is characterized by: S4, tracking and detecting the dynamic monomer to obtain the motion trajectory of the dynamic monomer, including: S41, establishing a motion model and an observation model of the tracking target, wherein the motion model describes the dynamic evolution process of the target state, and the observation model describes the relationship between the target state and the observation quantity; wherein the target state includes the position and speed of the target, and the observation quantity is the target detection result in continuous image frames; S42, predicting a priori estimated value of the target state of the current frame based on the target state estimated value of the previous frame and the motion model; Calculate the target state likelihood estimate of the current frame based on the target state prior estimate and the observation model of the current frame; S43, updating the target state posterior estimate of the current frame through Bayesian reasoning according to the target state prior estimate and the target state likelihood estimate as the tracking output; S44, calculating the similarity between the target state estimation values of the previous and next frames, and if the similarity is greater than a threshold, connecting the two target state estimation values into the same trajectory to obtain the motion trajectory of the dynamic monomer.
8. The electronic cemetery and information-based cemetery management method according to any one of claims 2 to 5, characterized in that: S6, constructing a virtual scene of the cemetery environment based on the 3D model of the static area and the motion trajectory of the dynamic monomer, including: S61, using a 3D rendering engine to load a 3D model of the static area to construct a static background of the virtual scene; S62, using the 3D human body model library to match the detection result of the dynamic monomer, instantiating the 3D human body model corresponding to the dynamic monomer, and rendering the human body model into the virtual scene according to the position of the dynamic monomer; S63, generating a motion animation of the human body model in the virtual scene according to the motion trajectory of the dynamic monomer; S64, using the scene graph management of the 3D rendering engine, performs fusion rendering of the static background and the dynamic human body model to generate a virtual scene of the cemetery environment.
9. The electronic cemetery and information-based cemetery management method according to claim 8 is characterized by: S63, generating a motion animation of the human body model in the virtual scene according to the motion trajectory of the dynamic monomer, including: Convert the position points on the dynamic monomer motion trajectory into the three-dimensional space coordinates of the human body model in the virtual scene coordinate system as the key frames of the human body model; According to the position change between two adjacent key frames of the human body model, the number of interpolation frames and the interpolation coefficient between adjacent key frames are calculated; Using a linear interpolation algorithm, a three-dimensional space coordinate sequence of an interpolation frame is generated according to the three-dimensional space coordinates and interpolation coefficients of adjacent key frames; Constructing a motion frame sequence of a human body model according to the three-dimensional space coordinate sequence of key frames and interpolation frames; In the three-dimensional rendering engine, the motion frame sequence of the human body model is loaded to control the movement of the human body model in the virtual scene.
10. An electronic cemetery and information-based mausoleum management system, characterized in that: include: At least one processing unit; used to execute instructions to implement the electronic cemetery and information-based cemetery management method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
A digital tomb-sweeping intelligent service system for cemeteries based on virtual reality
CN117974394B