A mine monitoring video playing method, device and electronic equipment

By updating the point cloud model in the 3D scene model of the mine and combining it with the pixel changes of the monitoring video, a 3D effect display of the mine monitoring video was achieved, which solved the problem of lack of 3D realism in traditional mine video monitoring technology and improved the monitoring effect and real-time performance.

CN120151499BActive Publication Date: 2025-12-23BEIJING AIENTROPY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510293202.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-12-23
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Traditional mine video surveillance technology lacks three-dimensional realism, cannot quickly locate the monitored area, and cannot display mining and operation conditions in real time.

Method used

By acquiring pixel changes in the monitoring video, the point cloud model in the 3D scene model of the mine is updated, and the 3D scene model is rendered in real time. Combined with the scenery around the point cloud model in the 3D scene model of the mine, a 3D effect display of the monitoring video is achieved.

Benefits of technology

It improves the playback effect of mine monitoring videos, enabling viewers to more accurately understand the location of the scenes captured by the monitoring videos, display the mining and operation situation in real time, and help monitoring personnel to discover safety hazards in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151499B_ABST
    Figure CN120151499B_ABST
Patent Text Reader

Abstract

The application discloses a mine monitoring video playing method and device and electronic equipment, relates to the digital twinborn technology field and the intelligent mine technology field, and comprises the following steps: acquiring pixel values of pixel points of a to-be-played video frame of a monitoring video, the monitoring video being obtained by shooting a specified area of a mine; determining each pixel point whose pixel value changes compared with a previous video frame in the to-be-played video frame; using the pixel values of the pixel points whose pixel values change to update pixel values of the same pixel points in a point cloud model loaded in a three-dimensional scene model of the mine, obtaining an updated three-dimensional scene model, the point cloud model being generated based on video frames of the monitoring video, and the point cloud model being loaded in a position corresponding to the specified area in the three-dimensional scene model; and rendering the updated three-dimensional scene model in real time. According to the scheme, the effect of playing the mine monitoring video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital twinning and the technical field of intelligent mines, and in particular to a mine monitoring video playing method and device and electronic equipment. BACKGROUND

[0002] With the continuous advancement of intelligent mine construction, the presentation mode of mine information is gradually changing from traditional two-dimensional plane to three-dimensional space. Usually, according to mine design CAD drawings, mine site videos, pictures and other related materials, a mine three-dimensional model is constructed, and a mine three-dimensional scene is built to show the overall situation of the mine and realize mine transparency.

[0003] With the development of network technology, more extensive video monitoring systems are built in mine areas to provide real-time monitoring pictures and assist mine personnel in responding quickly in emergency situations.

[0004] As one of the important means of mine production supervision, the role of video monitoring cannot be ignored. However, the traditional mine video monitoring technology often uses a flat list visualization method, which lacks a three-dimensional sense of reality and cannot quickly locate the monitoring area. In addition, although a three-dimensional scene of the mine can be built by placing a three-dimensional model to show the overall situation of the mine, it cannot show the mining and operation situation of the mine in real time. SUMMARY

[0005] The embodiments of the present application provide a mine monitoring video playing method and device and electronic equipment to solve the problem of poor mine monitoring video playing effect in the prior art.

[0006] The embodiments of the present application provide a mine monitoring video playing method, which comprises:

[0007] Obtaining pixel values of pixel points of a to-be-played video frame of a monitoring video, the monitoring video being obtained by shooting a specified area of a mine;

[0008] Determining each pixel point in the to-be-played video frame whose pixel value changes compared with a previous video frame;

[0009] Using the pixel values of each pixel point whose pixel value changes to update the pixel value of a same pixel point in a point cloud model loaded in a three-dimensional scene model of the mine, to obtain an updated three-dimensional scene model, the point cloud model being generated based on video frames of the monitoring video and loaded in a position corresponding to the specified area in the three-dimensional scene model;

[0010] Real-time rendering the updated three-dimensional scene model.

[0011] Further, the point cloud model is generated by the following steps:

[0012] obtaining a video key frame of a monitoring video collected in a specified area of a mine, the video key frame being a video frame in a sequence of continuous video frames that can more represent a scene of the specified area than other video frames;

[0013] identifying a three-dimensional scene graph of the mine and image feature points of the video key frame;

[0014] performing feature point matching according to attribute values of the three-dimensional scene graph and the image feature points of the video key frame;

[0015] assigning three-dimensional coordinates of the image feature points in the three-dimensional scene graph to the matched image feature points in the video key frame;

[0016] generating three-dimensional coordinates of each pixel point of the video key frame based on the three-dimensional coordinates of the image feature points of the video key frame that have been assigned with the three-dimensional coordinates;

[0017] generating a point cloud model according to the three-dimensional coordinates and pixel values of each pixel point of the video key frame.

[0018] Further, after the step of generating a point cloud model according to the three-dimensional coordinates and pixel values of each pixel point of the video key frame, the method further comprises:

[0019] loading the point cloud model into a corresponding position of the three-dimensional scene model according to the three-dimensional coordinates of each pixel point of the point cloud model, to obtain a three-dimensional scene model loaded with the point cloud model.

[0020] Further, the step of identifying a three-dimensional scene graph of the mine and image feature points of the video key frame comprises:

[0021] identifying the three-dimensional scene graph of the mine and the image feature points of the video key frame by using a SIFT algorithm, to obtain a feature point list of the three-dimensional scene graph and a feature point list of the video key frame;

[0022] The feature point list contains indexes and descriptor vectors of the image feature points.

[0023] Further, the step of performing feature point matching according to attribute values of the three-dimensional scene graph and the image feature points of the video key frame comprises:

[0024] calculating distances of descriptor vectors between each image feature point of the three-dimensional scene graph and each image feature point of the video key frame as descriptor distances, based on the indexes and descriptor vectors of the image feature points contained in the feature point list of the three-dimensional scene graph and the feature point list of the video key frame;

[0025] According to a matching strategy that a smaller descriptor distance indicates a higher matching degree of two image feature points, determine the matching image feature points between the three-dimensional scene graph and the video key frame.

[0026] Further, the three-dimensional coordinates of the image feature points with the three-dimensional coordinates assigned based on the video key frame are used to generate the three-dimensional coordinates of each pixel point of the video key frame, including:

[0027] The three-dimensional coordinates of the image feature points with the three-dimensional coordinates assigned based on the video key frame are used to generate the three-dimensional coordinates of each pixel point of the video key frame by using a coordinate interpolation algorithm.

[0028] Further, the pixel values of the pixel points of the to-be-played video frame of the monitoring video are obtained, including:

[0029] The pixel values of the pixel points of the to-be-played video frame at the same time of a plurality of monitoring videos are obtained, the plurality of monitoring videos being obtained by photographing a plurality of specified regions of the mine respectively;

[0030] The pixel values of the pixel points of the to-be-played video frame at the same time of a plurality of monitoring videos are obtained, the plurality of monitoring videos being obtained by photographing a plurality of specified regions of the mine respectively;

[0031] For each monitoring video, the pixel values of the pixel points of the to-be-played video frame of the monitoring video are used to update the pixel values of the same pixel points in the point cloud model corresponding to the monitoring video loaded in the three-dimensional scene model of the mine, the point cloud model corresponding to the monitoring video being generated based on the video frames of the monitoring video.

[0032] The embodiment of the application also provides a mine monitoring video playing device, including:

[0033] A video frame acquisition module is configured to obtain the pixel values of the pixel points of a to-be-played video frame of a monitoring video, the monitoring video being obtained by photographing a specified region of a mine;

[0034] A pixel value comparison module is configured to determine each pixel point in the to-be-played video frame whose pixel value has changed compared with a previous video frame;

[0035] A point cloud model updating module is configured to use the pixel values of the pixel points whose pixel values have changed to update the pixel values of the same pixel points in a point cloud model loaded in a three-dimensional scene model of the mine, to obtain an updated three-dimensional scene model, the point cloud model being generated based on the video frames of the monitoring video, and the point cloud model being loaded in a position corresponding to the specified region in the three-dimensional scene model.

[0036] A model rendering module is configured to render the updated three-dimensional scene model in real time.

[0037] Further, further comprising:

[0038] A key frame acquisition module is configured to acquire a video key frame of a monitoring video collected for a specified area of a mine, the video key frame being a video frame in a sequence of continuous video frames that can more represent a scene of the specified area than other video frames;

[0039] A feature point identification module is configured to identify image feature points of a three-dimensional scene graph of the mine and the video key frame;

[0040] A feature point matching module is configured to perform feature point matching according to attribute values of the image feature points of the three-dimensional scene graph and the video key frame;

[0041] A coordinate assignment module is configured to assign three-dimensional coordinates of the image feature points in the three-dimensional scene graph to matched image feature points in the video key frame;

[0042] A coordinate generation module is configured to generate three-dimensional coordinates of each pixel point of the video key frame based on the three-dimensional coordinates of the image feature points of the video key frame to which the three-dimensional coordinates are assigned;

[0043] A point cloud model generation module is configured to generate a point cloud model according to the three-dimensional coordinates and pixel values of each pixel point of the video key frame.

[0044] Further, further comprising:

[0045] A model loading module is configured to load the point cloud model to a corresponding position of the three-dimensional scene model according to the three-dimensional coordinates of each pixel point of the point cloud model, to obtain a three-dimensional scene model loaded with the point cloud model.

[0046] Further, the feature point identification module is specifically configured to identify the image feature points of the three-dimensional scene graph of the mine and the video key frame by using a SIFT algorithm, to obtain a feature point list of the three-dimensional scene graph and a feature point list of the video key frame;

[0047] The feature point list contains indexes and descriptor vectors of the image feature points.

[0048] Further, the feature point matching module is specifically configured to calculate distances of descriptor vectors between each image feature point of the three-dimensional scene graph and each image feature point of the video key frame as descriptor distances, based on the indexes and descriptor vectors of the image feature points contained in the feature point list of the three-dimensional scene graph and the feature point list of the video key frame;

[0049] According to a matching strategy that a smaller descriptor distance indicates a higher matching degree of two image feature points, the matching image feature points between the three-dimensional scene graph and the video key frame are determined.

[0050] Further, the coordinate generation module is specifically configured to generate three-dimensional coordinates of each pixel point of the video key frame by using a coordinate interpolation algorithm based on the three-dimensional coordinates of the image feature points of the video key frame to which the three-dimensional coordinates are assigned.

[0051] Further, the video frame acquisition module is specifically configured to acquire pixel values of pixel points of a to-be-played video frame at the same time of a plurality of monitoring videos, the plurality of monitoring videos being obtained by photographing a plurality of specified regions of the mine respectively.

[0052] Further, the point cloud model updating module is specifically configured to, for each monitoring video, update pixel values of the same pixel points in a point cloud model corresponding to the monitoring video loaded in the three-dimensional scene model of the mine by using the pixel values of the pixel points of the to-be-played video frame of the monitoring video, the point cloud model corresponding to the monitoring video being generated based on the video frames of the monitoring video.

[0053] The embodiment of the present application further provides an electronic device, including a processor and a machine readable storage medium, the machine readable storage medium stores machine executable instructions which can be executed by the processor, and the processor is prompted by the machine executable instructions to implement the mine monitoring video playing method.

[0054] The embodiment of the present application further provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the mine monitoring video playing method.

[0055] The embodiment of the present application further provides a computer program product containing instructions, when the computer program product is executed on a computer, the computer program product makes the computer execute the mine monitoring video playing method.

[0056] The beneficial effects of the present application include:

[0057] In the method provided by the embodiment of the present application, for the to-be-played video frame of the monitoring video to be displayed, each pixel point whose pixel value changes compared with the pixel value of the previous video frame in the to-be-played video frame is determined based on the pixel value of the pixel point, and the pixel value of the same pixel point in the point cloud model loaded in the three-dimensional scene model of the mine is updated using the pixel value of each pixel point whose pixel value changes, to obtain an updated three-dimensional scene model, and the updated three-dimensional scene model is rendered in real time. In the method, the point cloud model is generated based on the video frame of the monitoring video, and the point cloud model can represent a three-dimensional effect, so that the point cloud model is loaded in the three-dimensional scene model, and the three-dimensional effect of the content of the to-be-played video frame can be displayed by rendering the three-dimensional scene model in real time, and the point cloud model is loaded in the position corresponding to the specified region in the three-dimensional scene model, so that the viewer can more accurately know the position of the scene photographed by the monitoring video in combination with the scene around the point cloud model in the three-dimensional scene model of the mine, that is, the effect of playing the monitoring video of the mine is improved.

[0058] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings. BRIEF DESCRIPTION OF DRAWINGS

[0059] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, illustrate embodiments of the present application, and are used to explain the present application, and do not constitute a limitation on the present application. In the drawings:

[0060] Figure 1 The flowchart of the mine monitoring video playing method provided by the embodiment of the present application is shown in the figure;

[0061] Figure 2 The schematic diagram of the three-dimensional scene model loaded with the point cloud model rendered in the embodiment of the present application is shown in the figure;

[0062] Figure 3 The flowchart of generating the point cloud model in the embodiment of the present application is shown in the figure;

[0063] Figure 4 The flowchart of generating the three-dimensional scene diagram of the mine and calculating the three-dimensional coordinates of the pixel points in the three-dimensional scene diagram in the embodiment of the present application is shown in the figure;

[0064] Figure 5 The structural schematic diagram of the mine monitoring video playing device provided by the embodiment of the present application is shown in the figure;

[0065] Figure 6 The structural schematic diagram of the mine monitoring video playing device provided by another embodiment of the present application is shown in the figure;

[0066] Figure 7 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0067] In order to give an implementation scheme for improving the effect of playing a mine monitoring video, an embodiment of the present application provides a mine monitoring video playing method, device and electronic device. The preferred embodiments of the present application are described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In the case of no conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0068] An embodiment of the present application provides a mine monitoring video playing method, as shown in Figure 1 The method comprises the following steps.

[0069] Step 11, obtaining pixel values of pixel points of a to-be-played video frame of a monitoring video, the monitoring video being obtained by shooting a specified area of a mine;

[0070] Step 12, determining each pixel point in the to-be-played video frame whose pixel value changes compared with a previous video frame;

[0071] Step 13, using the pixel values of each pixel point whose pixel value changes to update the pixel value of a same pixel point in a point cloud model loaded in a three-dimensional scene model of the mine, to obtain an updated three-dimensional scene model, the point cloud model being generated based on a video frame of the monitoring video, and the point cloud model being loaded in a position corresponding to the specified area in the three-dimensional scene model;

[0072] Step 14, real-time rendering the updated three-dimensional scene model.

[0073] By using the above mine monitoring video playing method provided by the present application, since the point cloud model is generated based on the video frame of the monitoring video, and the point cloud model can show a three-dimensional effect, the point cloud model is loaded in the three-dimensional scene model, and the three-dimensional scene model is real-time rendered, so that the three-dimensional effect of the content of the to-be-played video frame can be displayed. In addition, the point cloud model is loaded in the position corresponding to the specified area in the three-dimensional scene model, so that the position of the scene shot by the monitoring video can be more accurately known in combination with the scene around the point cloud model in the three-dimensional scene model of the mine, that is, the effect of playing the mine monitoring video is improved.

[0074] In an embodiment of the present application, a three-dimensional scene model can be created in advance for the actual scene of the mine, and the three-dimensional scene model can display the three-dimensional effect of the mine scene.

[0075] The cameras can be arranged at multiple positions in the actual scene of the mine to respectively capture monitoring videos, and the cameras correspond one-to-one to the specified areas captured, that is, the monitoring videos captured correspond one-to-one to the specified areas.

[0076] For the captured monitoring videos, a point cloud model representing the scene of the specified area corresponding to the monitoring video can also be generated based on the video frames of the monitoring video, and the point cloud model is loaded into the position corresponding to the specified area in the three-dimensional scene model.

[0077] When there are multiple monitoring videos, and correspondingly, multiple point cloud models are generated, the multiple point cloud models can be loaded into the corresponding positions in the three-dimensional scene model.

[0078] In an embodiment of the present application, for the above step 11, the pixel values of the pixel points of the to-be-played video frames at the same time of the multiple monitoring videos can be obtained, the multiple monitoring videos are obtained by respectively capturing multiple specified areas of the mine, and the multiple to-be-played video frames obtained from the multiple monitoring videos are all captured at the same time;

[0079] Correspondingly, for the above step 13, for each monitoring video, the pixel values of the pixel points of the to-be-played video frames of the monitoring video that change can be used to update the pixel values of the same pixel points in the point cloud model corresponding to the monitoring video loaded in the three-dimensional scene model of the mine, and the point cloud model corresponding to the monitoring video is generated based on the video frames of the monitoring video;

[0080] After updating the point cloud model, the obtained three-dimensional scene model loads multiple latest point cloud models, and the multiple latest point cloud models can represent the content of the multiple to-be-played video frames, so that the synchronous playback of the multiple to-be-played video frames is realized by real-time rendering of the updated three-dimensional scene model.

[0081] And in the real-time rendered three-dimensional scene model, the video frames of the multiple monitoring videos are all located in the respective corresponding specified areas, which, combined with the scenery around the point cloud model in the three-dimensional scene model of the mine, can enable the viewer to more accurately know the positions of the scenes captured by the monitoring videos.

[0082] Figure 2 For the above mine monitoring video playback method provided by the embodiments of the present application, a schematic diagram of the rendered three-dimensional scene model loaded with the point cloud model is shown, wherein one monitoring video is taken as an example.

[0083] In an embodiment of the present application, the following method flow is proposed for generating a point cloud model, as shown in Figure 3 The method flow includes the following steps:

[0084] Step 31, obtain a video key frame of a monitoring video collected in a specified area of a mine, the video key frame being a video frame in a continuous video frame sequence that can more represent a scene of the specified area than other video frames.

[0085] A video is a sequence of continuous images, usually played at a certain frame rate, showing a dynamic visual effect, and the standard frame rate is 30 FPS (frames per second), that is, 30 images are contained in one second, and this frame rate is sufficient to capture smooth video. A video key frame is a representative frame in a video sequence that can help users quickly locate key scenes in the video. Usually, a frame image with relatively stable motion or little scene change and containing the main objects and activities in the scene is selected for subsequent processing.

[0086] In order to improve the efficiency of subsequent feature point matching, in this step, a video frame that maintains consistency with the scene objects contained in the specified area of the three-dimensional scene graph of the mine as much as possible can be selected as a video key frame.

[0087] In the embodiment of the application, the video key frame can be selected by a person, and in this step, the video key frame selected by the person is directly obtained, and subsequent operations are performed.

[0088] Step 32, identify image feature points of the three-dimensional scene graph of the mine and the video key frame.

[0089] In this step, specifically, the SIFT (Scale-Invariant Feature Transform) algorithm can be used to identify image feature points of the three-dimensional scene graph of the mine and the video key frame, to obtain a feature point list of the three-dimensional scene graph and a feature point list of the video key frame, and the feature point list contains indexes and descriptor vectors of the image feature points.

[0090] Further, the feature point list can further contain the following attribute information of the image feature points:

[0091] Feature point coordinates pt: a tuple of (x, y), where x is the horizontal coordinate and y is the vertical coordinate, with the upper left corner of the image as the origin, the x-axis to the right, and the y-axis downward;

[0092] Feature point response intensity response: representing the degree to which the point is a feature point, quantifying the significance or prominence of the feature point relative to its field;

[0093] Feature point field diameter size: in the neighborhood with a diameter of size, the response intensity of the feature point is the strongest.

[0094] Feature point descriptor des: namely descriptor vector, in SIFT algorithm, each feature point has a 128-dimensional descriptor vector, taking the feature point as the center, the feature point neighborhood is divided into 16*16 pixel regions, and 4*4 sub-regions are divided, the sub-regions are not overlapped, 8 direction gradient histograms are calculated in each sub-region (each direction is spaced by 45 degrees, namely 8-dimensional feature vector), and finally all 8-dimensional feature vectors are spliced to obtain a 128-dimensional descriptor vector, which is used for subsequent feature point matching by using Flann (Fast Library for Approximate Nearest Neighbors, fast search of nearest neighbors) algorithm.

[0095] Step 33, according to the attribute values of the image feature points of the three-dimensional scene graph and the video key frame, performing feature point matching.

[0096] In this step, the distance between the descriptor vectors of each image feature point of the three-dimensional scene graph and each image feature point of the video key frame can be calculated based on the indexes and descriptor vectors of the image feature points contained in the feature point list of the three-dimensional scene graph and the feature point list of the video key frame, as the descriptor distance.

[0097] According to the matching strategy that the smaller the descriptor distance is, the greater the matching degree of two image feature points is, the matching image feature points between the three-dimensional scene graph and the video key frame are determined.

[0098] Specifically, Flann algorithm can be used to read the three-dimensional scene graph, the feature point list of the three-dimensional scene graph, the video key frame, and the feature point list of the video key frame for feature point matching.

[0099] The feature point list of the three-dimensional scene graph and the feature point list of the video key frame are traversed, and the distance between the descriptor vectors of each feature point is calculated as the descriptor distance according to the above feature point descriptor des, if the ratio of the closest distance to the second closest distance is less than a threshold value (usually 0.7 to 0.8), it can be determined that this is a better match, and finally the matching relationship list of the feature points in the three-dimensional scene graph and the video key frame is obtained, and the matching relationship object mainly contains three attributes queryIdx, trainIdx and distance:

[0100] queryIdx: the index of the feature point of the three-dimensional scene graph in the feature point list of the three-dimensional scene graph;

[0101] trainIdx: the index of the matched feature point in the video key frame in the feature point list of the video key frame;

[0102] distance: The distance between feature points in the 3D scene graph and feature points in the video keyframe. The smaller the value, the greater the matching degree between the two feature points.

[0103] Furthermore, these attributes can be used to draw feature points and feature point matching in 3D scene graphs and video keyframes to obtain matching images, which can then be manually checked to ensure that the feature point matching is used correctly.

[0104] During the manual inspection of feature point matching, the matching status of feature points can be viewed based on the matching images described above. When there is an incorrect match, the corresponding feature point in the feature point list is deleted according to the index in the matching relationship, and the matching process is repeated to ensure the accuracy of feature point matching.

[0105] Step 34: Assign the 3D coordinates of the image feature points in the 3D scene image to the matching image feature points in the video keyframes.

[0106] In this step, the 3D coordinates of pixels in the 3D scene image can be assigned to the matching pixels in the video keyframes based on the obtained feature point matching relationship.

[0107] Step 35: Based on the three-dimensional coordinates of the image feature points with assigned three-dimensional coordinates in the video keyframe, generate the three-dimensional coordinates of the other pixels in the video keyframe.

[0108] In this step, specifically, based on the three-dimensional coordinates of the image feature points assigned with three-dimensional coordinates in the video keyframe, a coordinate interpolation algorithm can be used to generate the three-dimensional coordinates of other pixels in the video keyframe.

[0109] This step uses a coordinate interpolation algorithm to fill in missing data points, thereby smoothing the data and improving its resolution. This ensures that all pixels in the video keyframes have three-dimensional coordinates, thus improving the display effect of the point cloud model loaded into the subsequent 3D mining scene model.

[0110] Step 36: Generate a point cloud model based on the 3D coordinates and pixel values ​​of each pixel in the video keyframes.

[0111] A point cloud model is a collection of a large number of three-dimensional coordinate points (X, Y, Z) used to represent the shape and structure of an object. Point cloud models can also include auxiliary information such as color, intensity, and normals.

[0112] In this step, we can use the PCL library (PCL is an open-source library for processing 3D point cloud data) to read video keyframes with 3D coordinates and generate a point cloud model. Each point in the point cloud model records its 3D coordinates (X, Y, Z), pixel value, and unique identifier (which can be used to update the pixel values ​​of the pixels in the point cloud model later).

[0113] Step 37, load the point cloud model to the corresponding position of the three-dimensional scene model according to the three-dimensional coordinates of each pixel point of the point cloud model, to obtain a three-dimensional scene model loaded with the point cloud model.

[0114] Since the video data of the monitoring video usually contains a large amount of redundant information, such as the similarity between consecutive frames, in the embodiment of the application, the accuracy of feature point matching is improved by extracting the key frames of the video, so that the generated point cloud model can better display the video content.

[0115] In the mine monitoring video playing method provided in the embodiment of the application, in the 3D rendering engine, according to the change of the pixel value of the pixel point in the monitoring video frame, the pixel value of the corresponding pixel point in the point cloud model is modified accordingly. Specifically, the following can be performed:

[0116] In the 3D rendering engine, when the scene needs to view the monitoring video, the video frames are traversed, the 3D rendering engine acquires each video frame, traverses the pixel points of the video frame, judges whether the pixel value of the pixel point of the previous frame changes with the pixel value of the pixel point of the current frame, determines which pixel points of the current frame change in pixel value, and then updates the pixel value of the pixel point in the corresponding point cloud model according to the unique identifier of the pixel point (this method can improve the calculation efficiency and improve the rendering efficiency of the 3D rendering engine). Finally, by loading the point cloud model and modifying the pixel value of the pixel point in the point cloud model, the traditional flat list loading method of the monitoring video is replaced to view the real-time monitoring video in the three-dimensional scene model of the mine, that is, the real-time fusion of the monitoring video and the three-dimensional scene model is realized, thereby improving the playing effect of the mine monitoring video.

[0117] In the embodiment of the application, the process shown in FIG. 4 can be used to generate the three-dimensional scene graph of the mine and calculate the three-dimensional coordinates of the pixel points in the three-dimensional scene graph, including the following steps: Figure 4

[0118] Step 41, make a three-dimensional scene model of the mine.

[0119] In this step, Blender software can be used to determine the layout, size, structure, color, material and texture of the model according to the collected CAD plan of the mine, on-site photos or videos of the mine and other materials, to construct the three-dimensional scene model of the mine (such as the three-dimensional model of the unloading station on the surface of the mine, the internal repair room of the roadway, etc.).

[0120] Step 42, build a three-dimensional scene and render a three-dimensional scene graph and a depth map.

[0121] ​Blender software can be used to adjust the placement position and orientation of the three-dimensional model according to the collected mine site video, photos and other related materials, combined with the real mine monitoring camera position, to build a real mine three-dimensional scene.

[0122] By configuring the light of the scene, the best lighting and perspective can be obtained, and the Blender software provides multiple rendering engines, such as Eevee (real-time rendering) and Cycles (physics-based rendering). Select the appropriate rendering engine, set the image resolution, sample number, select the "depth" channel and other rendering parameters, and render the three-dimensional scene graph and depth map.

[0123] The depth map is a special image that contains the distance information from each pixel point in the scene to the camera, which can be used for subsequent three-dimensional coordinate calculation of the pixel points in the three-dimensional scene graph.

[0124] Step 43, calculate the three-dimensional coordinates of the pixel points in the three-dimensional scene graph.

[0125] Considering that the three-dimensional scene graph and the depth map are different rendering images of the same three-dimensional scene, the pixel points in the two images correspond one-to-one, and the pixel points in the depth map record the distance to the camera, therefore the three-dimensional coordinates of each pixel point in the depth map can be calculated, and the three-dimensional coordinates of the pixel points in the three-dimensional scene graph can be obtained. The following steps are used to obtain the three-dimensional coordinates of the pixel points in the depth map combined with the camera-related parameters:

[0126] First, calculate the coordinates of the pixel points in the normalized device (NDC, Normalized Device Coordinates) coordinate system, and the NDC coordinate range is from -1 to 1, that is, convert the pixel coordinates of the pixel points in the depth map to normalized coordinates (NDC coordinates).

[0127] In this step, the normalized coordinates can be calculated using the following formula:

[0128]

[0129] Where (u, v) is the pixel coordinate, and from the top left corner, horizontally to the right is u, and vertically downward is v, (sx, sy) is the width and height of the sensor size, (x n ,y n ) is the normalized coordinate.

[0130] Then, convert the normalized coordinates to camera coordinates in the camera coordinate system.

[0131] In this step, the normalized coordinates can be converted to camera coordinates using the following formula:

[0132] X c =d(u,v)·x n ;

[0133] Y c =d(u,v)·y n ;

[0134] Z c =d(u,v);

[0135] wherein d(u,v) is a depth value of the pixel point (u,v), i.e. a distance from the camera to the observed point, (X c ,Y c ,Z c ) is a camera coordinate.

[0136] In this step, it is assumed that the origin of the camera coordinate system is the camera position, and the Z axis points to the front of the camera.

[0137] Then, the camera coordinates are converted into world coordinates in the world coordinate system, and the world coordinates are three-dimensional coordinates.

[0138] In this step, the camera coordinates can be converted into world coordinates by using the following formula:

[0139]

[0140] wherein R is a rotation matrix in the camera parameters, T is a translation vector in the camera parameters, (X w ,Y w ,Z w ) is a world coordinate.

[0141] By using the above mine monitoring video playing method provided by the embodiments of the present application, the fusion of the monitoring video and the three-dimensional scene of the mine can be realized in real time, so that the mining and operation conditions of the mine are more vividly and intuitively displayed, the potential safety hazards are found in time by the monitoring personnel, the production process is better understood by the management personnel, the production strategy is timely adjusted, and the production efficiency is improved.

[0142] Based on the same inventive concept, according to the mine monitoring video playing method provided by the above embodiments of the present application, correspondingly, another embodiment of the present application further provides a mine monitoring video playing device, a structure diagram of which is shown in Figure 5 , and specifically includes:

[0143] The video frame acquisition module 501 is configured to acquire pixel values of pixel points of a to-be-played video frame of a monitoring video, wherein the monitoring video is obtained by photographing a specified area of a mine;

[0144] The pixel value comparison module 502 is configured to determine each pixel point in the to-be-played video frame whose pixel value changes compared with a previous video frame;

[0145] The point cloud model updating module 503 is configured to update pixel values of the same pixel points in a point cloud model loaded in a three-dimensional scene model of the mine by using pixel values of the pixel points whose pixel values change, to obtain an updated three-dimensional scene model, wherein the point cloud model is generated based on video frames of the monitoring video, and the point cloud model is loaded in a position corresponding to the specified area in the three-dimensional scene model.

[0146] The model rendering module 504 is configured to render the updated three-dimensional scene model in real time.

[0147] Further, as shown in Figure 6 , the method further comprises:

[0148] The key frame obtaining module 505 is configured to obtain a video key frame of the monitoring video collected for the specified area of the mine, wherein the video key frame is a video frame in a sequence of continuous video frames and is more capable of representing a scene of the specified area than other video frames.

[0149] The feature point identifying module 506 is configured to identify image feature points of a three-dimensional scene graph of the mine and the video key frame.

[0150] The feature point matching module 507 is configured to perform feature point matching according to attribute values of the image feature points of the three-dimensional scene graph and the video key frame.

[0151] The coordinate assigning module 508 is configured to assign three-dimensional coordinates of the image feature points in the three-dimensional scene graph to the matching image feature points in the video key frame.

[0152] The coordinate generating module 509 is configured to generate three-dimensional coordinates of each pixel point of the video key frame based on the three-dimensional coordinates of the image feature points of the video key frame to which the three-dimensional coordinates are assigned.

[0153] The point cloud model generating module 510 is configured to generate a point cloud model according to the three-dimensional coordinates and pixel values of each pixel point of the video key frame.

[0154] Further, as shown in Figure 6 , the method further comprises:

[0155] The model loading module 511 is configured to load the point cloud model into a corresponding position of the three-dimensional scene model according to the three-dimensional coordinates of each pixel point of the point cloud model, to obtain a three-dimensional scene model loaded with the point cloud model.

[0156] Further, the feature point identifying module 506 is specifically configured to identify the image feature points of the three-dimensional scene graph of the mine and the video key frame by using a SIFT algorithm, to obtain a feature point list of the three-dimensional scene graph and a feature point list of the video key frame.

[0157] The feature point list includes indexes and descriptor vectors of image feature points.

[0158] Further, the feature point matching module 507 is specifically configured to calculate distances of descriptor vectors between each image feature point of the three-dimensional scene graph and each image feature point of the video key frame as descriptor distances based on indexes and descriptor vectors of image feature points included in the feature point list of the three-dimensional scene graph and the feature point list of the video key frame.

[0159] According to a matching strategy that a smaller descriptor distance represents a higher matching degree of two image feature points, the matching image feature points between the three-dimensional scene graph and the video key frame are determined.

[0160] Further, the coordinate generation module 509 is specifically configured to generate three-dimensional coordinates of each pixel point of the video key frame by using a coordinate interpolation algorithm based on the three-dimensional coordinates of the image feature points of the video key frame to which the three-dimensional coordinates are assigned.

[0161] Further, the video frame acquisition module 501 is specifically configured to acquire pixel values of pixel points of a to-be-played video frame at the same time of a plurality of monitoring videos, the plurality of monitoring videos being obtained by photographing a plurality of specified regions of a mine respectively.

[0162] Further, the point cloud model updating module 503 is specifically configured to, for each monitoring video, update pixel values of the same pixel points in a point cloud model corresponding to the monitoring video in the three-dimensional scene model of the mine by using pixel values of each pixel point of the to-be-played video frame of the monitoring video, the point cloud model corresponding to the monitoring video being generated based on video frames of the monitoring video.

[0163] The functions of the above modules can correspond to the functions of the respective processing steps in the flowchart shown in the figure, and will not be described here again. Figures 1 to 4 The functions of the above modules can correspond to the functions of the respective processing steps in the flowchart shown in the figure, and will not be described here again.

[0164] The mine monitoring video playing device provided by the embodiment of the present application can be implemented by a computer program. Those skilled in the art should understand that the above-mentioned module division manner is only one of many module division manners, and as long as the mine monitoring video playing device has the above-mentioned functions, the division into other modules or without module division should be within the protection scope of the present application.

[0165] The embodiment of the present application also provides an electronic device, such as a mobile phone, a tablet computer, a personal computer, a server, etc. Figure 7As shown, the electronic device includes a processor 71 and a machine readable storage medium 72 storing machine executable instructions executable by the processor 71, the machine executable instructions causing the processor 71 to implement any of the above described mine monitoring video playing methods.

[0166] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement any of the above described mine monitoring video playing methods.

[0167] The embodiments of the present application also provide a computer program product containing instructions, which, when executed on a computer, cause the computer to perform any of the above described mine monitoring video playing methods.

[0168] The machine readable storage medium in the above described electronic device can include a random access memory (RAM) and can also include a non-volatile memory (NVM), for example, at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.

[0169] The processor described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0170] Each of the embodiments in the specification is described in a related manner, and the same or similar parts between each of the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the device, the electronic device, the computer readable storage medium and the computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0171] It is to be understood that the terms "including", "comprising", or any other variation thereof, are intended to cover the contents "open", such that a process, a method, an article, or an apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or even inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the stated element.

[0172] The present application is described with reference to the flowchart illustrations and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present application. It is understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, an embedded processor or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 The flowchart illustrations and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present application. Figure 1 The flowchart illustrations and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present application.

[0173] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device that implements the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 The flowchart illustrations and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present application. Figure 1 The flowchart illustrations and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present application.

[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 The flowchart illustrations and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present application. Figure 1 The flowchart illustrations and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present application.

[0175] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A mine monitoring video play method, characterized by, The method comprises the following steps: acquiring pixel values of pixel points of a to-be-played video frame of a monitoring video, the monitoring video being obtained by photographing a specified area of a mine; determining each pixel point in the to-be-played video frame whose pixel value has changed compared with a previous video frame; updating pixel values of the same pixel points in a point cloud model loaded in a three-dimensional scene model of the mine by using the pixel values of each pixel point whose pixel value has changed, to obtain an updated three-dimensional scene model, the point cloud model being generated based on video frames of the monitoring video, and the point cloud model being loaded in a position corresponding to the specified area in the three-dimensional scene model; real-time rendering the updated three-dimensional scene model; generating the point cloud model by using the following steps: acquiring a video key frame of the monitoring video collected in the specified area of the mine, the video key frame being a video frame in a continuous video frame sequence which can more represent a scene of the specified area than other video frames; identifying a three-dimensional scene graph of the mine and image feature points of the video key frame; performing feature point matching according to attribute values of the three-dimensional scene graph and the image feature points of the video key frame; assigning three-dimensional coordinates of the image feature points in the three-dimensional scene graph to the matched image feature points in the video key frame; generating three-dimensional coordinates of each pixel point of the video key frame based on the three-dimensional coordinates of the image feature points of the video key frame which have been assigned with the three-dimensional coordinates; generating a point cloud model according to the three-dimensional coordinates and pixel values of each pixel point of the video key frame.

2. The method of claim 1, wherein, After the step of generating the point cloud model according to the three-dimensional coordinates and pixel values of each pixel point of the video key frame, the method further comprises the following steps: loading the point cloud model into a corresponding position of the three-dimensional scene model according to the three-dimensional coordinates of each pixel point of the point cloud model, to obtain a three-dimensional scene model loaded with the point cloud model.

3. The method of claim 1, wherein, The step of identifying the three-dimensional scene graph of the mine and the image feature points of the video key frame comprises the following steps: identifying the three-dimensional scene graph of the mine and the image feature points of the video key frame by using a SIFT algorithm, to obtain a feature point list of the three-dimensional scene graph and a feature point list of the video key frame; the feature point list containing indexes and descriptor vectors of the image feature points.

4. The method of claim 3, wherein, The step of performing feature point matching according to attribute values of the three-dimensional scene graph and the image feature points of the video key frame comprises the following steps: calculating distances of descriptor vectors between each image feature point of the three-dimensional scene graph and each image feature point of the video key frame as descriptor distances, based on the indexes and descriptor vectors of the image feature points contained in the feature point list of the three-dimensional scene graph and the feature point list of the video key frame; determining the image feature points matched between the three-dimensional scene graph and the video key frame according to a matching strategy that a smaller descriptor distance represents a higher matching degree of two image feature points.

5. The method of claim 1, wherein, The step of generating three-dimensional coordinates of each pixel point of the video key frame based on the three-dimensional coordinates of the image feature points of the video key frame which have been assigned with the three-dimensional coordinates comprises the following steps: Based on the three-dimensional coordinates of the image feature points of the video key frame to which the three-dimensional coordinates are assigned, a coordinate interpolation algorithm is used to generate the three-dimensional coordinates of each pixel point of the video key frame.

6. The method of claim 1, wherein, The pixel value of each pixel point of the to-be-played video frame of the monitoring video is obtained, including: The pixel values of the pixel points of the to-be-played video frame at the same time of a plurality of monitoring videos are obtained, the plurality of monitoring videos being obtained by photographing a plurality of specified regions of the mine respectively; The pixel value of each pixel point of the to-be-played video frame of the monitoring video is obtained, including: For each monitoring video, the pixel value of each pixel point of the to-be-played video frame of the monitoring video is used to update the pixel value of the same pixel point in the point cloud model loaded in the three-dimensional scene model of the mine corresponding to the monitoring video, the point cloud model corresponding to the monitoring video being generated based on the video frames of the monitoring video.

7. A mine monitoring video playback device, characterized by, Including: A video frame acquisition module is configured to obtain the pixel value of each pixel point of the to-be-played video frame of the monitoring video, the monitoring video being obtained by photographing a specified region of the mine; A pixel value comparison module is configured to determine each pixel point of the to-be-played video frame whose pixel value has changed compared with that of a previous video frame; A point cloud model updating module is configured to use the pixel value of each pixel point whose pixel value has changed to update the pixel value of the same pixel point in the point cloud model loaded in the three-dimensional scene model of the mine, to obtain an updated three-dimensional scene model, the point cloud model being generated based on the video frames of the monitoring video, and the point cloud model being loaded in a position corresponding to the specified region in the three-dimensional scene model; A model rendering module is configured to render the updated three-dimensional scene model in real time. Further, the method further includes: A key frame acquisition module is configured to obtain a video key frame of the monitoring video collected for the specified region of the mine, the video key frame being a video frame in a continuous video frame sequence that can more represent the scene of the specified region than other video frames; A feature point identification module is configured to identify image feature points of the video key frame and a three-dimensional scene map of the mine; A feature point matching module is configured to perform feature point matching according to attribute values of the image feature points of the three-dimensional scene map and the video key frame; A coordinate assignment module is configured to assign the three-dimensional coordinates of the image feature points in the three-dimensional scene map to the matched image feature points in the video key frame; A coordinate generation module is configured to generate the three-dimensional coordinates of each pixel point of the video key frame based on the three-dimensional coordinates of the image feature points of the video key frame to which the three-dimensional coordinates are assigned; A point cloud model generation module is configured to generate a point cloud model according to the three-dimensional coordinates and pixel values of each pixel point of the video key frame.

8. An electronic device, comprising: The machine readable storage medium stores machine executable instructions which can be executed by the processor, and the processor is prompted by the machine executable instructions to implement the method of any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method in any one of claims 1-6.

Citation Information

Patent Citations

  • A construction method of a three-dimensional GIS dynamic model

    CN109598794A

  • Construction method and system of multi-temporal live-action three-dimensional model and terminal equipment

    CN116797744A