Asset management method, device and storage medium based on model and video fusion
By building a twin model of physical assets and virtual assets, combined with video fusion technology, the problem of time-consuming and labor-consuming manual inventory is solved, real-time monitoring and prediction of park assets is achieved, and corporate management efficiency and trust are improved.
Patent Information
- Application Number
- CN202510712285.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-30
AI Technical Summary
In the prior art, asset management relies on manual regular inventory, which is time-consuming and labor-intensive and lacks attention to digital assets.
By building a physical asset twin model and a virtual asset twin model, combining video fusion technology, the virtual and real mapping of the device and multi-dimensional data fusion are realized, and real-time monitoring and prediction are carried out.
It has achieved all-round protection of park assets, ensured the authenticity and integrity of data, improved the level of enterprise supervision and resource allocation efficiency, and improved the performance of enterprise.
Smart Images

Figure CN120235711B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of logistics warehousing asset management, and in particular to an asset management method, device and storage medium based on model and video fusion. Background Art
[0002] Assets can be divided into different categories based on relevant standards and attributes. Assets can be both physical assets and non-physical assets, such as digital assets. Depending on the type of asset, corresponding asset management strategies can be formulated to better manage and protect assets.
[0003] It's well-known that enterprise asset management contributes to corporate value creation, and efficient asset management can significantly improve business performance. In logistics and warehousing applications, changes in fixed and current assets are the most common physical assets, such as increases or decreases in machinery, inventory, and raw materials. The most traditional asset management method relies on regular manual inventory checks, but this approach is both time-consuming and labor-intensive, and lacks attention to digital assets. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an asset management method, device and storage medium based on model and video fusion, so as to solve the problem in the prior art that the asset management method relies on manual regular inventory, which is time-consuming and labor-intensive, and lacks attention to digital assets.
[0005] According to a first aspect of an embodiment of the present invention, there is provided an asset management method based on model and video fusion, comprising:
[0006] Obtaining the park's warehousing and logistics asset data, and dividing the park's warehousing and logistics asset data into two categories: physical asset data and digital asset data;
[0007] Constructing a virtual asset twin model based on the operational data and knowledge model in the digital asset data;
[0008] Based on the devices included in the physical asset data, select the 3D model of each device from a pre-built local model library, build a virtual campus environment through digital twins, and place the 3D model of each device at a corresponding location in the virtual campus environment;
[0009] Video images of each device are acquired through real cameras set up at different locations in the park. The UV texture coordinates of the effective graphics elements of each device's 3D model are calculated based on the geometric mapping relationship, and a dynamic mapping relationship between the effective graphics elements of each device and the corresponding video images is established.
[0010] By extracting the video image of each device frame by frame and refreshing the surface texture buffer data of the 3D model of each device according to the UV texture coordinates of the valid graphics elements, the surface texture of the 3D model of each device is made consistent with the video image, and a twin model of the physical asset is constructed.
[0011] Preferably, it also includes:
[0012] Establishing a twin high-definition camera digital model in the virtual campus environment that matches the spatial parameters of a real camera, wherein the shooting angle and viewing range of the twin high-definition camera digital model are consistent with those of the real camera;
[0013] Obtain a frame of image from the fixed-view video data shot by the real camera, extract key feature information of the image through a feature point detection algorithm, and calculate the initial pose of the real camera;
[0014] Images taken from the same perspective by a real camera and the corresponding twin high-definition camera digital model are obtained respectively. The bundle adjustment algorithm is used to iteratively optimize the reprojection error between the real space points and the virtual projection points to obtain high-precision camera pose parameters after nonlinear optimization. The initial pose of the real camera is adjusted using the high-precision camera pose parameters to achieve real camera calibration.
[0015] Preferably,
[0016] Calculating the UV texture coordinates of the valid primitives of the 3D model of each device based on the geometric mapping relationship includes:
[0017] Obtaining the homogeneous coordinates of any vertex of a valid primitive of the device's 3D model in world space, obtaining a rotation matrix for rotating the vertex from the world coordinate system to the camera coordinate system, and obtaining a displacement matrix for translating the vertex from the world coordinate system to the camera coordinate system; and obtaining the homogeneous coordinates of any vertex in camera space based on the homogeneous coordinates of the arbitrary vertex in world space, the rotation matrix, and the displacement matrix;
[0018] Get the focal length of the real camera, and based on the projection relationship of any vertex from camera space to image space, use the homogeneous coordinates of any vertex in camera space and the focal length of the real camera to get the coordinates of any vertex in image space;
[0019] Get the image width and height, and get the texture coordinates of any vertex in the texture space based on the image width, height and the coordinates of any vertex in the image space;
[0020] The texture coordinates of each vertex of the valid primitive of the 3D model of the device are obtained, and the UV texture coordinates of the valid primitive of the 3D model of each device are obtained.
[0021] Preferably, it also includes:
[0022] Obtain the functional attributes of the physical entities of each 3D model in the physical asset twin model, bind each 3D model with the functional attributes of the corresponding physical entity, realize the virtual-reality mapping between the 3D model and the physical entity, and dynamically update the functional attributes of the physical entity of each 3D model in real time through the monitoring data of the physical entity of each 3D model.
[0023] Preferably, it also includes:
[0024] Based on the multi-dimensional data generated by each device in the virtual asset twin model during operation, a pre-built multi-level deep fusion perception model is used to predict the health classification results of each device.
[0025] Preferably, it also includes:
[0026] The health classification results of each device predicted by using the pre-built multi-level deep fusion perception model include:
[0027] Each dimension of the multi-dimensional data generated by the device during operation is regarded as a mode, and the multiple modes of the device are input into independent encoder networks respectively, and each mode is encoded into a feature vector separately;
[0028] The feature vector of each modality is mapped to the attention weight space through a fully connected layer, and then the softmax function is used to calculate the attention weight of each feature vector;
[0029] Perform weighted summation on each feature vector according to the corresponding attention weight to obtain the fusion representation;
[0030] The fused representation is passed through a fully connected layer to generate a model output, which is the health classification result of the device.
[0031] Preferably, it also includes:
[0032] The degradation trend is obtained based on the health classification results of each device, and a preset number of health classification results and corresponding degradation trends are input into a pre-built time series prediction model. The pre-built time series prediction model outputs the health status of the device in the future.
[0033] According to a second aspect of an embodiment of the present invention, there is provided an asset management device based on model and video fusion, comprising:
[0034] Data acquisition module: used to obtain the park's warehousing and logistics asset data, and divide the park's warehousing and logistics asset data into two categories: physical asset data and digital asset data;
[0035] Virtual asset mapping module: used to build a virtual asset twin model based on the operating data and knowledge model in the digital asset data;
[0036] Entity model and environment mapping module: used to select the 3D model of each device from the pre-built local model library according to the devices included in the physical asset data, build a virtual campus environment through digital twins, and place the 3D model of each device at the corresponding position in the virtual campus environment;
[0037] Video fusion module: used to obtain video images of each device through real cameras set up at different locations in the park; calculate the UV texture coordinates of the effective graphics elements of each device's 3D model based on the geometric mapping relationship, and establish a dynamic mapping relationship between the effective graphics elements of each device and the corresponding video image;
[0038] Physical asset mapping module: used to extract the video image of each device frame by frame and refresh the surface texture buffer data of the 3D model of each device according to the UV texture coordinates of the valid graphics element, so that the surface texture of the 3D model of each device is consistent with the video image, and build a physical asset twin model.
[0039] According to a third aspect of an embodiment of the present invention, a storage medium is provided, which stores a computer program. When the computer program is executed by a main controller, it implements each step of the logistics equipment redesign method based on digital twins.
[0040] The technical solutions provided by the embodiments of the present invention may have the following beneficial effects:
[0041] This application obtains the park's warehousing and logistics asset data and divides the park's warehousing and logistics asset data into two categories: physical asset data and digital asset data; constructs twin models of each park's warehousing and logistics asset, which are divided into physical asset twin models and virtual asset twin models; among them, when constructing the physical asset twin model, the digital twin technology is combined with video splicing and fusion technology to realize the virtual-real mapping of the physical entity of the equipment and the 3D model, and realize the time-space consistency fusion of the real video stream and the virtual three-dimensional scene; through the solution of this application, all-round protection of the park's physical assets and digital assets is achieved, ensuring the authenticity and integrity of the data, increasing the trust between various links in supply chain management, improving the level of corporate supervision, optimizing resource allocation, and ultimately achieving the goal of improving corporate performance and corporate value.
[0042] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0044] Figure 1 is a flowchart illustrating an asset management method based on model and video fusion according to an exemplary embodiment;
[0045] Figure 2 is a system diagram of an asset management device based on model and video fusion according to another exemplary embodiment;
[0046] In the attached figure: 1-data acquisition module, 2-virtual asset mapping module, 3-physical model and environment mapping module, 4-video fusion module, 5-physical asset mapping module. DETAILED DESCRIPTION
[0047] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.
[0048] Example 1
[0049] Figure 1 is a flow chart of an asset management method based on model and video fusion according to an exemplary embodiment. Figure 1 As shown, the method includes:
[0050] S1, obtaining the park's warehousing and logistics asset data, and dividing the park's warehousing and logistics asset data into two categories: physical asset data and digital asset data;
[0051] S2, constructing a virtual asset twin model based on the operating data and knowledge model in the digital asset data;
[0052] S3, based on the devices included in the physical asset data, selecting a 3D model of each device from a pre-built local model library, building a virtual campus environment through digital twins, and placing the 3D model of each device at a corresponding location in the virtual campus environment;
[0053] S4, using real cameras set up at different locations in the park to obtain video images of each device; calculating the UV texture coordinates of the effective graphics elements of the 3D model of each device based on the geometric mapping relationship, and establishing a dynamic mapping relationship between the effective graphics elements of each device and the corresponding video image;
[0054] S5, extracting the video image of each device frame by frame and refreshing the surface texture buffer data of the 3D model of each device according to the UV texture coordinates of the valid primitives, so that the surface texture of the 3D model of each device is consistent with the video image, and constructing a physical asset twin model;
[0055] It is understandable that this embodiment first obtains and collects the park's warehousing and logistics asset data, and classifies the data obtained, dividing all disposable assets in the park into two categories: physical asset data and digital asset data. The collected data includes but is not limited to the architectural drawings and structural designs of the entire park, warehousing and logistics equipment data, park operation data, equipment operation and maintenance fault data, park personnel data, park security data, park financial data, park energy consumption data, environmental parameters, etc.
[0056] Secondly, twin models of various assets in the park's warehousing and logistics are constructed, which are divided into physical asset twin models and virtual asset twin models. The physical asset twin model refers to the physical storage equipment and park environment in the park, which can be expressed through model entities. The virtual asset twin model is the data and knowledge model generated in the park. For the physical asset twin model, it can be provided by the pre-built local model library. It is worth noting that the 3D model provided by the local model library is only the overall structure and outline of the equipment, and does not include the surface texture details of the equipment. Therefore, this application will subsequently update the surface texture of the model through video fusion. The 3D model can also be constructed and uploaded in real time through 3D scanning. When using the pre-built model library solution, the common general equipment models in the park are stored in the MinIO model library, and the model is called through the model library. MinIO is a high The high-performance, lightweight object storage server is designed for large-scale data storage and analysis. In addition to data storage, it can also implement data encryption, providing a foundation for subsequent data management of model digitization. At the same time, it converts and encrypts the 3D device models in the MinIO model library, further reducing the load when calling models in the model library. The conversion process converts the 3D device model into code and stores it on the server. For non-standard models, they can also be uploaded to the MinIO local model library through real-time 3D point cloud scanning or drawing with model building software. When the model is called through the MinIO model library, for some specific or special physical devices, it is not necessary to perform complete physical modeling. The model outline is outlined according to the approximate size of the physical device, and a simple model is established. The mapping of the physical device is achieved through later video fusion and splicing technology.
[0057] Video fusion and stitching technology: A certain number of high-definition cameras are deployed in the park to monitor the entire park and specific areas. Video fusion and stitching technology is to fuse the images captured by the cameras with the constructed twin model, so that the captured images are mapped to the model, that is, the video images are mapped to the model as a texture map. This embodiment must ensure both the range captured by the camera and the clarity of the captured images and videos so that they are not distorted during the fusion and mapping process. At the same time, the video mapping results must be adjusted based on certain strategies to adapt to changes in the model under multi-view observation. Since the images captured by the real camera are only captured from one or several angles, in order to perfectly map the real-time images captured to the twin model to achieve virtual-real fusion, it is necessary to implement real-scene video coordinate calibration and transformation, specifically including:
[0058] Image acquisition: Images from the same perspective are acquired from a real camera and its corresponding digital twin model (virtual camera). The digital twin model simulates the imaging characteristics of the real camera through parameter settings (such as focal length and sensor size) to ensure that the perspectives of the two are consistent.
[0059] Reprojection error modeling: There is a geometric deviation between the spatial points captured by the real camera and the projected points of the virtual model, which is called reprojection error. The bundle adjustment algorithm minimizes this error through nonlinear optimization. Specifically, it adjusts the camera pose (such as rotation and translation parameters) and spatial point coordinates to make the projected virtual points coincide with the real points as much as possible.
[0060] Iterative optimization process: The Levenberg-Marquardt algorithm is used to iteratively optimize the reprojection error and solve the least squares solution of the objective function. The optimization variables include the camera pose parameters (extrinsic parameters) and the 3D point coordinates. The Jacobian matrix of the parameters is updated and iterated by calculating the error, gradually converging to the optimal solution.
[0061] Pose parameter adjustment: The optimized high-precision pose parameters are used to correct the initial pose of the real camera, eliminating calibration errors caused by installation deviations or environmental noise, and ultimately achieving accurate camera calibration. When the camera pose changes, only the above calibration process needs to be repeated.
[0062] After the camera pose calibration is completed, the following 3D model fusion operations need to be performed: First, the UV texture coordinates of the valid primitives in the 3D model (the primitive is the simplest geometric shape that constitutes the 3D model, and the default here is a triangle) are calculated based on the geometric mapping relationship, and a dynamic mapping relationship between the primitives and the video image is established; then, the input source of the texture sampler is configured to the real-time monitoring video stream. By extracting the video image frame by frame and refreshing the texture buffer data, the surface texture of the 3D model and the video image are kept synchronized and updated in real time, ultimately achieving the spatiotemporal consistency of the real video stream and the virtual 3D scene. Specifically,
[0063] To render the captured images and videos as textures, we first need to calculate the texture coordinates of each vertex of all the primitives in the 3D model. The UV texture coordinates of the valid primitives are calculated as follows: The primitive vertex is calculated in the image camera Taking the texture coordinate process as an example, if the homogeneous coordinates of the vertex in world space are First, we need to transform the vertex from world space to camera space. Assume that the rotation matrix and displacement matrix from world space to camera space are respectively and , then the homogeneous coordinates of the vertex in camera space are , where the vertex is the homogeneous coordinate in world space, because It is defined in the world coordinate system, so it is a known quantity; The rotation matrix describes the rotation relationship between the world coordinate system and the camera coordinate system. It can be determined by camera calibration or known camera posture and is a known quantity. The displacement matrix describes the translation relationship between the world coordinate system and the camera coordinate system. It is used to represent the position of the camera in the world coordinate system. It can also be determined by camera calibration or known camera position. It is a known quantity. Therefore, the homogeneous coordinates of the unknown vertex in the camera space can be calculated according to the formula ;
[0064] To transform the vertex from camera space to image space, if the vertex coordinates in camera space are (is a known quantity, represents the transpose of the matrix), and then linearly transforms it to the image space coordinates on the imaging plane as , for this formula we have is the focal length of the camera, and is the coordinate on the image plane, in pixels. According to the principle of triangle similarity, we can get , which represents the projection relationship from the camera coordinate system to the image coordinate system, and further simplification gives , we can get the image space coordinates ;
[0065] Finally, the vertices need to be transformed from image space to texture space. Since the origin of the image space coordinate system is at the center of the image, and the origin of the texture space coordinate system is at the lower left corner, it is necessary to normalize the coordinate system, that is, divide the horizontal and vertical coordinates by the width and height of the corresponding image. Assuming the image width is , the height is , then the final texture coordinates and image space coordinates The relationship is , and finally the texture coordinates can be obtained ;
[0066] After completing the construction of the twin model for the physical asset, it is necessary to bind the model attributes to map the model to the physical entity. The functional attributes of the asset are defined through attributes such as location, time, temperature, and status. These attributes are used to represent the basic characteristics and functions of the asset and represent the status information of the asset. Each asset object initializes these attribute values according to actual needs. During the asset life cycle, these attribute values can be dynamically updated according to actual monitoring data, realizing real-time reflection of the asset status.
[0067] The virtual asset twin model integrates multi-dimensional data generated during the operation of physical equipment (such as operating data, fault records, maintenance records, etc.), extracts core knowledge that represents the equipment's attributes and status, and builds a data-driven intelligent analysis module based on this. It applies deep learning, large models and other technologies to deeply mine the equipment's full life cycle data, builds health evaluation standards based on the virtual asset model, realizes real-time diagnosis and trend prediction of equipment health status, and forms a structured and iterative equipment knowledge system. Specifically, it includes:
[0068] As assets continue to accumulate, a database of equipment will be formed, and based on the database, a knowledge base of equipment will be formed. Therefore, based on the mastery of a large amount of data and knowledge, a data- and knowledge-driven predictive modeling method is proposed - a multi-level deep fusion perception model based on data-knowledge feature decision-making. Because the data collected in this application is multi-dimensional, multimodal fusion technology can perceive and understand multimodal data from different sources, and improve the comprehensive understanding of the environment through cross-modal learning and association, thereby establishing connections between different data sources. In addition, it also has the ability to make adaptive decisions and plans based on environmental changes and information updates;
[0069] This example assumes that there are multiple modal inputs , are processed by independent encoder networks, and the input of each modality is encoded separately as a feature vector , that is, ,in For the The encoder network of the modality.
[0070] The eigenvector of each mode All pass through a fully connected layer Mapped to the attention weight space, and then use the softmax function to calculate its relative attention weight , which is , where the softmax function ensures that the sum of the weights of all modes is 1, so that the contribution of different modes to the final output can be dynamically adjusted;
[0071] After obtaining the attention weight of each modality, the feature vectors of each modality are Perform weighted summation according to their corresponding weights to generate a fusion representation , that is, ,This step realizes the information integration between different modalities,,enabling the model to capture the complementary information in,multimodal data;
[0072] Fusion Representation Further processing through the fully connected layer finally generates the output of the model , which is ,in and are the weight matrix and bias vector of the fully connected layer respectively. Finally, Predict classification results for integrated multi-source data;
[0073] Based on the multi-source prediction and classification results, a health evaluation standard system is established, which covers two dimensions: real-time status diagnosis and long-term trend prediction;
[0074] Real-time status diagnosis: The value sequence is evaluated, and thresholds are set for different equipment types, industry standards, expert knowledge or historical data to classify equipment into clear different status levels (such as healthy, sub-healthy, fault (different fault types, etc.)). The classification probability vector value directly corresponds to different state levels.
[0075] Long-term trend forecast: Based on the real-time status diagnosis results, Extracting the degradation trend, when accumulating a certain amount of history After obtaining the value and the corresponding degradation trend, it is used as input to predict its future health assessment curve based on the time series prediction model (such as LSTM).
[0076] Example 2
[0077] Figure 2 1 is a system diagram of an asset management device based on model and video fusion according to another exemplary embodiment, the device comprising:
[0078] Data acquisition module 1: used to obtain the park's warehousing and logistics asset data, and divide the park's warehousing and logistics asset data into two categories: physical asset data and digital asset data;
[0079] Virtual asset mapping module 2: used to build a virtual asset twin model based on the operating data and knowledge model in the digital asset data;
[0080] Entity model and environment mapping module 3: used to select the 3D model of each device from the pre-built local model library according to the devices included in the physical asset data, build a virtual campus environment through digital twins, and place the 3D model of each device at the corresponding position in the virtual campus environment;
[0081] Video Fusion Module 4: Used to acquire video images of each device using real cameras installed at different locations in the park; calculate the UV texture coordinates of the effective graphics elements of each device's 3D model based on the geometric mapping relationship, and establish a dynamic mapping relationship between the effective graphics elements of each device and the corresponding video image;
[0082] Physical asset mapping module 5: It is used to extract the video image of each device frame by frame and refresh the surface texture buffer data of the 3D model of each device according to the UV texture coordinates of the valid graphics element, so that the surface texture of the 3D model of each device is consistent with the video image, and build a physical asset twin model.
[0083] Example 3:
[0084] This embodiment provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a host controller, each step in the above method is implemented;
[0085] It is understandable that the storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0086] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.
[0087] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" is at least two.
[0088] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0089] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0090] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0091] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0092] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0093] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0094] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. The asset management method based on model and video fusion is characterized by: include: Obtaining the park's warehousing and logistics asset data, and dividing the park's warehousing and logistics asset data into two categories: physical asset data and digital asset data; Constructing a virtual asset twin model based on the operational data and knowledge model in the digital asset data; Based on the devices included in the physical asset data, select the 3D model of each device from a pre-built local model library, build a virtual campus environment through digital twins, and place the 3D model of each device at a corresponding location in the virtual campus environment; The video images of each device are acquired through real cameras set up at different locations in the park; Calculate the UV texture coordinates of the effective graphics elements of the 3D model of each device based on the geometric mapping relationship, and establish a dynamic mapping relationship between the effective graphics elements of each device and the corresponding video image; Calculating the UV texture coordinates of the valid primitives of the 3D model of each device based on the geometric mapping relationship includes: Get the homogeneous coordinates of any vertex of a valid primitive in the device's 3D model in world space, get the rotation matrix that rotates the vertex from the world coordinate system to the camera coordinate system, and get the displacement matrix that translates the vertex from the world coordinate system to the camera coordinate system; Obtaining the homogeneous coordinates of any vertex in camera space according to the homogeneous coordinates of any vertex in world space, the rotation matrix, and the displacement matrix; Get the focal length of the real camera, and based on the projection relationship of any vertex from camera space to image space, use the homogeneous coordinates of any vertex in camera space and the focal length of the real camera to get the coordinates of any vertex in image space; Get the image width and height, and get the texture coordinates of any vertex in the texture space based on the image width, height and the coordinates of any vertex in the image space; Obtain the texture coordinates of each vertex of the valid primitive of the 3D model of the device, and obtain the UV texture coordinates of the valid primitive of the 3D model of each device; By extracting the video image of each device frame by frame and refreshing the surface texture buffer data of the 3D model of each device according to the UV texture coordinates of the valid graphics elements, the surface texture of the 3D model of each device is made consistent with the video image, and a twin model of the physical asset is constructed.
2. The method according to claim 1, characterized in that Also includes: Establishing a twin high-definition camera digital model in the virtual campus environment that matches the spatial parameters of a real camera, wherein the shooting angle and viewing range of the twin high-definition camera digital model are consistent with those of the real camera; Obtain a frame of image from the fixed-view video data shot by the real camera, extract key feature information of the image through a feature point detection algorithm, and calculate the initial pose of the real camera; Images taken from the same perspective by a real camera and the corresponding twin high-definition camera digital model are obtained respectively. The bundle adjustment algorithm is used to iteratively optimize the reprojection error between the real space points and the virtual projection points to obtain high-precision camera pose parameters after nonlinear optimization. The initial pose of the real camera is adjusted using the high-precision camera pose parameters to achieve real camera calibration.
3. The method according to claim 2, characterized in that Also includes: Obtain the functional attributes of the physical entities of each 3D model in the physical asset twin model, bind each 3D model with the functional attributes of the corresponding physical entity, realize the virtual-reality mapping between the 3D model and the physical entity, and dynamically update the functional attributes of the physical entity of each 3D model in real time through the monitoring data of the physical entity of each 3D model.
4. The method according to claim 3, characterized in that Also includes: Based on the multi-dimensional data generated by each device in the virtual asset twin model during operation, a pre-built multi-level deep fusion perception model is used to predict the health classification results of each device.
5. The method according to claim 4, characterized in that Also includes: The health classification results of each device predicted by using the pre-built multi-level deep fusion perception model include: Each dimension of the multi-dimensional data generated by the device during operation is regarded as a mode, and the multiple modes of the device are input into independent encoder networks respectively, and each mode is encoded into a feature vector separately; The feature vector of each modality is mapped to the attention weight space through a fully connected layer, and then the softmax function is used to calculate the attention weight of each feature vector; Perform weighted summation on each feature vector according to the corresponding attention weight to obtain the fusion representation; The fused representation is passed through a fully connected layer to generate a model output, which is the health classification result of the device.
6. The method according to claim 5, characterized in that Also includes: The degradation trend is obtained based on the health classification results of each device, and a preset number of health classification results and corresponding degradation trends are input into a pre-built time series prediction model. The pre-built time series prediction model outputs the health status of the device in the future.
7. Asset management device based on model and video fusion, characterized in that: include: Data acquisition module: used to obtain the park's warehousing and logistics asset data, and divide the park's warehousing and logistics asset data into two categories: physical asset data and digital asset data; Virtual asset mapping module: used to build a virtual asset twin model based on the operating data and knowledge model in the digital asset data; Entity model and environment mapping module: used to select the 3D model of each device from the pre-built local model library according to the devices included in the physical asset data, build a virtual campus environment through digital twins, and place the 3D model of each device at the corresponding position in the virtual campus environment; Video fusion module: used to obtain video images of each device through real cameras set up at different locations in the park; Calculate the UV texture coordinates of the effective graphics elements of the 3D model of each device based on the geometric mapping relationship, and establish a dynamic mapping relationship between the effective graphics elements of each device and the corresponding video image; Calculating the UV texture coordinates of the valid primitives of the 3D model of each device based on the geometric mapping relationship includes: Obtaining the homogeneous coordinates of any vertex of a valid primitive of the device's 3D model in world space, obtaining a rotation matrix for rotating the vertex from the world coordinate system to the camera coordinate system, and obtaining a displacement matrix for translating the vertex from the world coordinate system to the camera coordinate system; and obtaining the homogeneous coordinates of any vertex in camera space based on the homogeneous coordinates of the arbitrary vertex in world space, the rotation matrix, and the displacement matrix; Get the focal length of the real camera, and based on the projection relationship of any vertex from camera space to image space, use the homogeneous coordinates of any vertex in camera space and the focal length of the real camera to get the coordinates of any vertex in image space; Get the image width and height, and get the texture coordinates of any vertex in the texture space based on the image width, height and the coordinates of any vertex in the image space; Obtain the texture coordinates of each vertex of the valid primitive of the 3D model of the device, and obtain the UV texture coordinates of the valid primitive of the 3D model of each device; Physical asset mapping module: used to extract the video image of each device frame by frame and refresh the surface texture buffer data of the 3D model of each device according to the UV texture coordinates of the valid graphics element, so that the surface texture of the 3D model of each device is consistent with the video image, and build a physical asset twin model.
8. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the main controller, each step of the asset management method based on model and video fusion as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Unified coding method and system for park equipment asset application Internet of Things acquisition
CN119697218A