Industrial material pile identification method, system and equipment and storage medium

By fusing radar point cloud and text description data into the pile type judgment model and combining it with environmental data to adjust features, the problem of insufficient high-level semantic modeling in pile identification is solved, high-accuracy and stable identification is achieved in complex environments, natural language feedback is supported, and the development of pile identification technology towards intelligent management and control is promoted.

CN120763701AActive Publication Date: 2025-10-10CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510903176.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-10
Estimated Expiration
2045-07-01

Smart Images

  • Figure CN120763701A_ABST
    Figure CN120763701A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, in particular to an industrial material pile identification method, system and device and a storage medium. The industrial material pile identification method comprises the following steps: obtaining to-be-identified material pile data; inputting the material pile data into a material pile type judgment model, processing the material pile data by the material pile type judgment model according to a preset rule, and outputting a material pile type identification result; wherein the processing steps of the material pile type judgment model on the material pile data are as follows: performing feature extraction on radar point cloud data to obtain point cloud features; performing feature extraction on the stacking state text description data to obtain semantic features; fusing the point cloud features and the semantic features to obtain fused features; analyzing the environment data to determine an environment influence weight, and adjusting the fusion feature according to the environment influence weight to obtain an environment perception fusion feature; and determining a material pile type identification result according to the environment perception fusion features. According to the invention, the accuracy of the material pile type identification result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an industrial material pile identification method, system, device and storage medium. Background Art

[0002] In industrial scenarios such as metallurgy, mining, and building materials, the dynamic management of material piles (such as ore, coal, sand and gravel, etc.) faces many challenges. First, if the parameters of the material pile such as the stacking height and tilt angle exceed the prescribed range, it may cause a serious collapse accident and bring huge safety hazards. Secondly, the traditional manual inspection method is difficult to achieve real-time acquisition of the three-dimensional morphological parameters of the material pile (such as volume, density distribution, etc.), which leads to efficiency bottlenecks. Especially in large-scale industrial production, it is difficult to ensure timely and accurate status feedback and it is difficult to form an automated industrial operation process. In addition, environmental interference factors at the industrial site, such as changes in light and dust obstruction, often greatly affect the data reliability of a single sensor, further exacerbating the difficulty of management and monitoring.

[0003] Currently, pile identification technology primarily relies on three approaches: point cloud-based, vision-based, and point cloud and image fusion. However, all three approaches have limitations, and a comprehensive framework for practical industrial application has yet to be established. The first category, point cloud-based processing, relies on 3D sensors such as radar or depth cameras and employs point cloud processing algorithms such as PointNet++ to capture the 3D geometric structure of the pile. While these approaches offer advantages in spatial modeling, they lack high-level semantic modeling capabilities and are unable to effectively identify semantic information such as material type and stacking status. The second category, vision-based approaches based on 2D images, typically employ convolutional neural networks (CNNs) for image recognition. While these approaches can identify the pile's appearance and surface texture, they are limited by camera viewing angles, lighting conditions, and occlusion, making it difficult to accurately estimate 3D metrics such as the pile's volume and spatial distribution. Fuel piles are often three-dimensional, with complex and diverse stacking configurations. Traditional image processing methods cannot fully represent these three-dimensional structures. 2D image processing techniques are particularly limited when processing point cloud data and complex spatial relationships. This method ignores the three-dimensional geometric information of the pile, making it difficult for the model to maintain high accuracy and stability in complex environments. The third category is the point cloud and image fusion solution, which attempts to combine the two types of modal information for pile state modeling through methods such as sensor registration and multi-channel feature splicing. Although it improves the recognition accuracy to a certain extent, it is limited by the accuracy of heterogeneous data alignment, the limited feature fusion strategy, and the lack of task-level semantic guidance, and its performance in industrial dynamic scenarios is still unstable. In addition, many existing pile type judgment methods rely on manual feature extraction and traditional machine learning algorithms, such as support vector machines (SVM) and decision trees. These methods usually require manual analysis and feature design of the pile shape. However, manual feature extraction is not only time-consuming but also difficult to capture complex patterns in the data. Especially when faced with large amounts of high-dimensional data, traditional machine learning methods are relatively insufficient in terms of accuracy and adaptability. As the amount of data increases, traditional methods often find it difficult to cope with more complex pile type classification tasks, limiting their effectiveness and efficiency in large-scale practical applications.

[0004] More critically, existing technologies generally lack the integration of Large Language Models (LLMs) into multimodal processing architectures, making them incapable of supporting natural language-based feedback and responses for interactive human-computer tasks. This is particularly true in complex industrial scenarios where comprehensive decision-making requires combining structural recognition with state understanding. The lack of interpretable and schedulable language modeling approaches severely restricts the expansion of stockpile identification technology into intelligent management and control. Summary of the Invention

[0005] The present application aims to at least solve the technical problems existing in the prior art and provide an industrial stockpile identification method, system, device and storage medium.

[0006] In a first aspect, the present invention provides an industrial stockpile identification method, comprising:

[0007] Obtaining data of the material pile to be identified, the material pile data including radar point cloud data of the material pile, text description data of the stacking status, and environmental data;

[0008] The pile data is input into the pile type judgment model, and the pile type judgment model processes the pile data according to preset rules and outputs the pile type recognition result;

[0009] The steps for the stockpile type judgment model to process the stockpile data are as follows:

[0010] Perform feature extraction on radar point cloud data to obtain point cloud features;

[0011] Perform feature extraction on the stacking status text description data to obtain semantic features;

[0012] Fuse point cloud features and semantic features to obtain fused features;

[0013] Analyze environmental data to determine environmental impact weights, and adjust fusion features based on the environmental impact weights to obtain environmental perception fusion features;

[0014] The pile type recognition result is determined based on the environmental perception fusion features.

[0015] In a second aspect, the present invention provides an industrial stockpile identification system, the system comprising:

[0016] An acquisition module is used to acquire the material pile data to be identified, the material pile data including radar point cloud data of the material pile, text description data of the stacking status and environmental data;

[0017] A processing module is used to input the stockpile data into the stockpile type judgment model, and the stockpile type judgment model processes the stockpile data according to preset rules and outputs the stockpile type recognition result;

[0018] The steps for the stockpile type judgment model to process the stockpile data are as follows:

[0019] Perform feature extraction on radar point cloud data to obtain point cloud features;

[0020] Perform feature extraction on the stacking status text description data to obtain semantic features;

[0021] Fuse point cloud features and semantic features to obtain fused features;

[0022] The environmental data is analyzed to determine an environmental influence weight, and a fusion feature is adjusted according to the environmental influence weight, to obtain an environmental perception fusion feature;

[0023] A stockpile type identification result is determined according to the environmental perception fusion feature.

[0024] In a third aspect, the present application provides an electronic device, which comprises:

[0025] at least one processor; and,

[0026] a memory connected to the at least one processor in communication; wherein,

[0027] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the industrial stockpile identification method described above.

[0028] In a fourth aspect, the present application further provides a computer readable storage medium, which stores at least one computer program, and the at least one computer program is executed by a processor in an electronic device to implement the industrial stockpile identification method described above.

[0029] In summary, the present application has the following beneficial technical effects:

[0030] The stockpile data is identified and judged by the stockpile type judgment model, the stockpile type judgment model comprehensively considers the point cloud feature, the semantic feature and the environmental data of the stockpile data, determines the environmental influence weight according to the environmental data, adjusts the fusion proportion of the point cloud feature and the semantic feature according to the environmental influence weight, the fusion feature includes the influence information of the environmental change on the stability of the stockpile, and the adaptability of the stockpile type judgment model in the complex changing environment can be enhanced, and the accuracy of the stockpile type judgment can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 A flowchart of an industrial stockpile identification method provided by an embodiment of the present application is shown;

[0032] Figure 2 A module diagram of a stockpile type judgment model provided by an embodiment of the present application is shown;

[0033] Figure 3 An internal processing flowchart of an environmental driving feature regulation module provided by an embodiment of the present application is shown;

[0034] Figure 4 A structural diagram of an electronic device for implementing the industrial stockpile identification method provided by an embodiment of the present application is shown.

[0035] Reference numerals: 10, processor; 11, memory; 12, communication bus; 13, communication interface.

[0036] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0037] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0038] In the description of the present invention, it should be understood that the terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention.

[0039] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal communication between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.

[0040] Reference Figure 1 FIG. 1 is a flow chart of an industrial material pile identification method according to an embodiment of the present invention. In this embodiment, the industrial material pile identification method includes:

[0041] S1. Obtain data of the material pile to be identified.

[0042] In industrial scenarios, industrial material piles include ore, coal, sand and gravel, metallurgical waste slag, fuel waste slag, glass waste slag, papermaking waste slag, and ceramic waste slag, etc. The material pile data includes radar point cloud data of the material pile, stacking status text description data and environmental data.

[0043] The radar point cloud data of the material pile is collected by a laser radar device (LiDAR), which can provide high-precision three-dimensional spatial point cloud data. Specifically, the point cloud data collected by the laser radar device is an N x 3 matrix, where N is the number of points, and each row represents the three-dimensional coordinate (x, y, z) data format of a point. In addition, the laser radar device may also be accompanied by intensity or point cloud color information, which helps to further analyze the quality of the point cloud data and the material characteristics of the object.

[0044] In order to remove irrelevant data and noise, a Voxel Grid Filter is used to downsample the point cloud data; this method can effectively reduce the amount of data and preserve the geometric shape features of the point cloud. After that, the downsampled point cloud data is standardized to obtain the final radar point cloud data. The specific method of standardizing the downsampled point cloud data is: the three-dimensional coordinates of each point are subtracted by the mean value and divided by the standard deviation to ensure that the data is within a unified scale, avoiding the influence of large-scale coordinate differences on subsequent processing. The final radar point cloud data is obtained through the point cloud data.

[0045] The stacking state text description data is used to represent the natural language description of the stacking state of the material pile; the stacking state text description data includes at least one of the type of the material pile, the stacking requirements, the environmental conditions, etc. The stacking state text description data can be derived from the standard documents of the material pile, the stacking requirements, the stacking safety specifications, etc.

[0046] After collecting the text description corresponding to the stacking state text description data, first, use a natural language processing toolkit (such as NLTK or SpaCy) to perform word segmentation on the collected text description, remove stop words (such as meaningless words such as "of", "and", etc.), and other preprocessing. The English full name of the natural language processing toolkit NLTK is Natural Language Toolkit, LTK is a widely used Python library that provides a wealth of natural language processing tools and resources, including part-of-speech tagging, syntax analysis, semantic analysis, etc. Its advantages are rich in functions, detailed documentation, and good community support; the natural language processing toolkit SpaCy is a powerful NLP library (Natural Language Processing), which provides high-performance natural language processing tools and models, including tokenization, word vector representation, named entity recognition, etc.

[0047] Then, the BERT (Bidirectional Encoder Representations from Transformers) pre-trained model is used to encode the preprocessed text description. The BERT model can generate fixed-dimensional vector representations that can capture the semantic information in the description as input for subsequent processing. After BERT processing, the embedding vector of the description data is a fixed-dimensional vector. In order to maintain the uniformity of the data, these fixed-dimensional vectors are L2 normalized to obtain the final stacked state text description data, ensuring that the text descriptions in the stacked state text description data are all on the same scale.

[0048] Environmental data includes at least one of temperature data, humidity data and vibration data. During the environmental data acquisition stage, environmental data of the material pile storage area is obtained in real time through multiple sensors (such as temperature sensors, humidity sensors, vibration sensors, etc.); the impact of the industrial material pile storage environment on the material pile shape is judged based on the data of the temperature sensor, humidity sensor and / or vibration sensor, and the impact of the environmental data on the material pile shape is fully considered in the subsequent identification of the material pile shape, so that the subsequent material pile shape identification results are more accurate.

[0049] S2. Input the material pile data into the material pile type judgment model. The material pile type judgment model processes the material pile data according to preset rules and outputs a material pile type recognition result.

[0050] In this embodiment, the pile shape determination model is an innovation and improvement based on the architecture of the GPT4Point model. The pile shape determination model is a multimodal learning model used for pile shape determination tasks. GPT4Point is a large multimodal point cloud model that can complete 3D object recognition, understanding, and question-answering tasks using only point clouds as input without the aid of images. It also combines point cloud understanding with generation, giving the large multimodal model the ability to control 3D generation.

[0051] Reference Figure 2 The material pile shape determination model consists of a point cloud feature extraction module, a descriptive feature extraction module, a multimodal fusion module, and an environmentally driven feature control module. The point cloud feature extraction module extracts geometric features from point cloud data using a Transformer network and a multi-scale convolutional network to generate point cloud features. The descriptive feature extraction module processes descriptive information using the BERT model to generate semantic features. The multimodal fusion module fuses point cloud features and semantic features to generate multimodal fusion features. This allows the model to comprehensively consider both point cloud features and feature descriptive information when determining material pile shape. Finally, the environmentally driven feature control module dynamically adjusts the weights of point cloud features by collecting environmental data to adapt to the impact of environmental changes on material pile determination.

[0052] Specifically, the pile type judgment model processes the pile data in the following steps:

[0053] S21. Extract features from the radar point cloud data to obtain point cloud features.

[0054] The point cloud feature extraction module mainly includes a multi-layer Transformer encoder and a multi-scale convolutional network. In the preferred implementation of this embodiment, the specific processing method of step S21 is:

[0055] S211. Input the point cloud data into a multi-layer Transformer encoder to obtain a geometric feature representation of the point cloud data.

[0056] In this embodiment, the point cloud feature extraction module uses an improved Transformer architecture, which can capture the relationship between points through a self-attention mechanism and process the geometric information of the pile. The input radar point cloud data has a dimension of N×3, where N is the number of points and 3 is the three-dimensional spatial coordinate (x, y, z) of each point. Based on this, a multi-layer Transformer encoder is used to process the point cloud data and generate a geometric feature representation. Specifically:

[0057] The radar point cloud data first passes through the Transformer encoder. Each layer of the Transformer processes the input point cloud data through the self-attention mechanism to capture the relationship between points. The calculation formula is as follows:

[0058] F Point =TransformerEncoder(X)

[0059] Among them, X is the input point cloud data, F Point is the feature representation of the point cloud after Transformer encoding; TransformerEncoder(·) represents the Transformer encoding operation.

[0060] S212. Use a multi-scale convolutional network to process the geometric feature representation to obtain local and global geometric features of the stockpile.

[0061] To further enhance point cloud feature extraction capabilities, a Multi-Scale Convolutional Neural Network (MS-CNN) was introduced. This network extracts both local and global geometric features of the stockpile through convolution operations at different scales. In a multi-scale convolutional network, local features are primarily extracted in shallow convolutional layers. Shallow convolutional layers convolve the input image using multiple convolution kernels, each of which extracts local geometric features in the image, such as edges and textures. Global features are primarily extracted in deeper convolutional layers. As the convolutional layers deepen, the local geometric features extracted in shallow layers are combined and abstracted into higher-level features, ultimately forming global geometric features.

[0062] The processing formula of the multi-scale convolutional network for geometric feature representation is as follows:

[0063] F i =Covn i (F Point ),i∈{1,2,3}

[0064] Among them, F i is the output of the i-th convolution operation, i represents the index of the convolution layer of the multi-scale convolutional network; F Point It is the radar point cloud feature after Transformer encoding.

[0065] S213: Determine the point cloud features of the stockpile based on the local geometric features and the global geometric features.

[0066] S22. Extract features from the stacking status text description data to obtain semantic features.

[0067] The stacking state text description data is processed by the BERT model, which generates high-quality embedding vectors that capture the characteristic information in the stacking state text description. After processing by the BERT model, each feature description is converted into a fixed-dimensional vector representation as input for subsequent processing. The core formula is as follows:

[0068] F text =BERT pooled (T) = CLS_token (Transformer L (T embed ))

[0069] Where T represents the input text description sequence, which consists of n tokens:

[0070] T=[t1,t2,…,t n ]

[0071] T embed=E token +E position +E segment

[0072] T embed is the embedding representation of the input text. Transformer L Represents the BERT backbone network consisting of L layers of Transformer encoders. CLS_token is the first special token (classification tag) output by BERT, which is used for overall semantic representation.

[0073] BERT here pooled (T) represents pooling of the BERT output. The most common approach is to extract the first CLS_token vector. This token is a special symbol BERT adds at the beginning of the input sequence to aggregate the semantic information of the entire text. It is particularly suitable for tasks that require holistic representation, such as text classification. Other strategies include average pooling, max pooling, and multi-layer feature fusion, but in this model, we use CLS_token as the default.

[0074] The input embedding layer can be expressed as

[0075] T embed [i]=E token (t i )+E position (i)+E segment (s i )

[0076] Among them E token (t i ) indicates token(t i )’s word vector embedding, E position (i) represents the position code of position i, using sine-cosine function:

[0077]

[0078] Among them, i represents the absolute position index of the token in the sequence, j represents the corresponding frequency order, and d model is the size of the hidden dimension of the embedding vector, and the constant 10000 is used to expand the frequency range on an exponential scale, thereby covering multi-scale position information from high to low. The exponential decay makes the sine wave corresponding to the first few dimensions have short periods (high frequency, sensitive to absolute / local position) and the latter few dimensions have long periods (low frequency, can encode global / relative distance). pos (i,2j) and E pos(i, 2j+1) correspond to the calculation of different dimensions of the position encoding vector respectively. j is used to traverse the dimension index and to determine whether the sine or cosine function should be used to calculate the different dimensions of the corresponding vector when calculating the position encoding.

[0079] E segment (s i ) represents segmentation code, which is used to distinguish different sentences.

[0080] BERT output sequence [C,h1,h2,……,h n ], where C = CLS_token.

[0081] C(the length of C is d model ) is the context embedding of the CLS tag at the beginning of the sequence, which is used for overall semantic representation. It contains the global semantic information of the entire input text (or text pair) and can be regarded as an aggregate expression of the overall meaning of the input text. n Each h (of the same length) in the 𝑠 corresponds to the contextual embedding of the i-th actual token. It is the output vector corresponding to each token in the input text after being processed by the BERT model. It contains the semantic information of the token at the corresponding position and also integrates the influence of other tokens in the context. It can be used for token-level tasks.

[0082] S23. Fuse the point cloud features and semantic features to obtain fused features.

[0083] In order to facilitate the fusion of point cloud features and semantic features, the semantic features are aligned with the point cloud features;

[0084] To align with the point cloud features, add a linear projection layer: F' text =LayerNorm(W proj F text +b proj )

[0085] in The BERT feature dimension d BERT Mapping to point cloud feature dimension d point .W proj is the weight matrix of the linear projection layer. It is a matrix with spatial dimensions The matrix is ​​responsible for converting the feature dimension d output by BERT BERT Mapping to point cloud feature dimension d point , so that the text features and point cloud features can be aligned in dimension, which facilitates subsequent operations such as fusion.

[0086] Point cloud features and semantic features are fused through a multimodal alignment mechanism. The alignment process uses a self-attention mechanism to calculate the similarity between point cloud features and description features, thereby achieving the fusion of the two and ultimately generating a unified feature representation. The alignment formula is as follows:

[0087] F aligned =Aliend(F Point ,F' text )

[0088] Among them, F aligned is the multimodal feature representation after fusion; F Point is the feature obtained by the point cloud coding network Transformer, F' text The descriptive text features are output by the text encoder BERT and enhanced to the same dimension through linear projection;

[0089] Aliend(·) is a multimodal alignment function. Aliend(·) uses cosine similarity as a metric to calculate and weight the fusion of two modal features to generate a unified cross-modal representation. aligned It not only preserves the spatial information of the point cloud but also incorporates the semantic information of the text. The alignment process is performed by calculating the similarity between the point cloud features and the point cloud description features. This patent uses the cosine similarity method.

[0090] First calculate the cosine similarity of point cloud features and semantic features:

[0091]

[0092] Among them, ||·|| means finding the modulus length;

[0093] Then perform weighted fusion:

[0094] F aligned =β·F Point +(1-β)·F' text

[0095] where β is expressed as

[0096] ∈ is a very small constant that prevents the denominator from being zero.

[0097] S24. Analyze the environmental data to determine the environmental impact weight, and adjust the fusion feature according to the environmental impact weight to obtain the environmental perception fusion feature.

[0098] In the preferred mode of the present embodiment, by collecting environmental data (such as temperature, humidity, vibration, etc.) in real time, automatically calculating the environmental impact weight, adjusting the fusion features according to the changes of environmental data, so that the model can maintain high accuracy and robustness in dynamic environment. The workflow of this module includes environmental data collection, influence function design, feature adjustment calculation and final weight application.

[0099] Referring to Figure 3 , the environmental data includes temperature data, humidity data and vibration data, wherein the temperature influence E temp on the environment is represented by the temperature of the area where the stockpile is located, and the temperature may affect the physical properties of the stockpile; the humidity influence E humidity on the environment is represented by the environmental humidity, which may change the stability of the stockpile and the material; the vibration influence E vibration on the environment is represented by the vibration sensor monitoring the vibration changes of the environment where the stockpile is located, and the vibration may cause the displacement or instability of the stockpile.

[0100] In the preferred embodiment of the present embodiment, the specific processing mode of step S23 is:

[0101] Determine the environmental impact weight according to at least one of the temperature influence factor, the humidity influence factor and the vibration influence factor, and determine the fusion ratio of the fusion feature and the semantic feature according to the environmental impact weight;

[0102] Determine the temperature influence factor of the environment on the stockpile type according to the temperature data, determine the humidity influence factor of the environment on the stockpile type according to the humidity data, and determine the vibration influence factor of the environment on the stockpile type according to the vibration data.

[0103] In the environment-driven feature regulation module, the core task is to calculate the influence of each environmental factor on the fusion feature according to the environmental data. This process is completed through the influence function . Each environmental data source (temperature, humidity, vibration) has a corresponding influence function, and through the output of these functions, the model can dynamically adjust the weight of the fusion feature. The present application designs different influence functions for each environmental data source, and the specific method is as follows:

[0104] The influence of temperature change on the stockpile is usually nonlinear, especially in high temperature environment, the stability of the stockpile may decrease. Therefore, the influence function of temperature needs to enhance the attention to the stability of the stockpile in the case of large temperature fluctuation. Specifically, the sigmoid function is used to represent the influence of temperature on the fusion feature. The sigmoid function can smoothly adjust the influence degree according to the range of temperature change, so as to appropriately weight the fusion feature. The calculation formula of temperature influence factor is

[0105]

[0106] wherein, represents the temperature influence factor; k temp is the sensitivity coefficient of temperature, used to represent the sensitivity of temperature change to the environmental influence weight; E temp represents the current temperature value, and T0 is the reference temperature value;

[0107] The influence of humidity on the pile is similar to that of temperature, and generally in extreme humidity conditions, the stability of the pile will change. The humidity influence function needs to appropriately increase the model's attention to the stability of the pile type in the case of large humidity changes. The ReLU (Rectified Linear Unit) function of the present application is used to represent the influence of humidity on the fused features. The ReLU function can quickly enhance the weight of the features when the humidity exceeds a certain threshold. The calculation formula of the humidity influence factor is

[0108]

[0109] wherein, represents the humidity influence factor; k humidity is the sensitivity coefficient of humidity, used to represent the influence of humidity on the environmental influence weight; E humidity is the current humidity value, and H0 is the reference humidity value; max(·) represents the maximum value operation;

[0110] The influence of vibration on the pile is very direct, especially in the environment with strong vibration, the stability of the pile will decrease significantly. Therefore, the influence function of vibration needs to enhance the influence on the pile type judgment when the vibration is strong. The present application adopts the combination form of sigmoid function and ReLU, so as to strongly affect the weight adjustment of the model in the case of large vibration amplitude. The calculation formula of the vibration influence factor is

[0111]

[0112] wherein, represents the vibration influence factor; k vibration is the sensitivity coefficient of vibration, E vibration is the current vibration amplitude, V0 is the reference vibration value, and V threshold is the vibration threshold value.

[0113] Each environment data source has a corresponding influence function, and the final environmental weight EnvWeigh t(E) is obtained by weighting and combining the influences of all environment data sources. The present application calculates the environmental weight by weighting and summing the influence functions of each environment data source, and the calculation formula of the environmental influence weight is

[0114]

[0115] Among them, EnvWeight(E) represents the environmental impact weight, w temp Indicates the weight of the temperature influence factor, w vibration Represents the weight of humidity impact factor, w humidit Indicates the weight of the vibration impact factor.

[0116] Specifically, the environmental impact weight EnvWeight(E) is used to dynamically adjust the fusion features to ensure that under different environmental conditions, the model can automatically enhance its focus on the most important features for determining the pile type. The adjustment formula for dynamically adjusting the original fusion weight of the point cloud features based on the environmental impact weight EnvWeight(E) is as follows:

[0117] F adjusted =F aligned ×EnvWeight(E)

[0118] Among them, F adjusted is the fusion feature after environmental perception adjustment (i.e., environmental perception fusion feature), F aligned is the fusion feature, and EnvWeight(E) is the weight calculated based on the environmental data.

[0119] S25. Determine the material pile type recognition result based on the environmental perception fusion feature.

[0120] The pile type determination model is tasked with accurately classifying and describing the pile type using fused features derived from feature extraction, fusion, and environmentally driven feature manipulation. Based on the fused environmentally perceived features, it generates a detailed description of the pile, including information such as the pile type, stability, and stacking pattern of the industrial pile.

[0121] The description of the pile shape is generated using the GPT-4 model (Generative Pre-trained Transformer 4), a pre-trained, large-scale generative Transformer model specifically designed for description generation tasks. GPT-4 uses input features to generate a natural language description of the pile shape. This process leverages GPT-4's powerful autoregressive generation capabilities to generate a detailed pile shape description from the feature vector in a step-by-step manner.

[0122] The input of the GPT-4 model is the fusion feature F adjusted , which is a feature vector fused through the manipulation of point cloud data, description information, and environmental features. This vector contains multi-dimensional information about the pile, including its geometry, semantic information, and environmental adaptability.

[0123] The GPT-4 model uses an autoregressive generation mechanism for description generation. The characteristic of the autoregressive model is that each time a new word is generated, it depends on the previously generated words. The generation process of the model can be expressed as:

[0124] y t =GPT-4(F adjusted ,y1,y2,…y t-1 )

[0125] Among them, y t is the output of the current time step (the tth word), F adjusted is the input environment perception fusion feature vector, which contains all necessary point clouds, descriptive features and environmental information, y1,y2,...,y t-1 is the first t-1 words generated. In this way, GPT-4 uses an autoregressive mechanism to rely on input features and previously generated words to gradually generate the entire description until it reaches a predetermined end marker or generates a complete description.

[0126] The generated description will accurately reflect the pile type, stacking method, stability, and environmental adaptability of the pile. The following is a specific example of the generated description:

[0127] (1) “The pile presents a three-dimensional stacking pattern, with a stable bottom but a risk of tilting at the top. It is suitable for stacking in a relatively stable environment.”

[0128] (2) “The pile is flat and has high overall stability, but may deform to some extent under high humidity conditions.”

[0129] (3) “The pile has a conical structure, which is suitable for long-term accumulation, but may be unstable in an environment with frequent vibrations.”

[0130] (4) “The piles are irregular in shape, scattered and stacked, and have poor overall stability. Safety monitoring of the stockpiling area needs to be strengthened.”

[0131] The final description output not only provides classification information of the stockpile, but also provides rich environmental adaptability information for subsequent management, storage and safety assessment.

[0132] In another embodiment of the present application, the industrial stockpile identification method further includes:

[0133] S3. Obtain a training data set, and use the training data set to train a stockpile type judgment model.

[0134] The training data set includes stockpile sample data and true information corresponding to the stockpile shape of the stockpile sample data. The training process of the stockpile shape judgment and description generation model aims to optimize the GPT-4 model so that it can generate accurate stockpile descriptions after receiving input features and capture key information such as shape type and stability during the generation process. The training process focuses on improving the description generation ability of the model by minimizing the description generation loss.

[0135] The training steps of the stockpile shape judgment model include:

[0136] S31, constructing the network structure of the initial stockpile shape judgment model.

[0137] S32, training the network of the initial stockpile shape judgment model using the training data set. In each training, the stockpile sample data is input into the initial stockpile shape judgment model, and the predicted result of the stockpile shape of the stockpile sample data by the initial stockpile shape judgment model is obtained.

[0138] The stockpile sample data includes point cloud data, description data and environment data. The point cloud data is processed by a feature extraction module to generate point cloud features, which are used to describe the geometric information of the stockpile. The description data is converted into description features by a BERT model, providing semantic support. Temperature, humidity, vibration and other environmental factors are smoothed by a Kalman filter to generate environmental impact weights.

[0139] The point cloud features, description features and environmental impact weights are input into the GPT-4 model, and the model is optimized by minimizing the generation loss. GPT-4 generates the description of the stockpile according to the input F adjusted features. The training goal is to optimize the description generation ability of the model by minimizing the generation loss.

[0140] S33, calculating the negative log-likelihood loss between the true information and the predicted result, and optimizing the network parameters of the initial stockpile shape judgment model according to the negative log-likelihood loss to obtain the final stockpile shape judgment model.

[0141] Since the task of stockpile shape judgment and description generation depends on GPT-4 to generate point cloud descriptions, the loss function mainly focuses on the quality of description generation. Therefore, the negative log-likelihood loss (NLLLoss) is used, which is used to measure the difference between the generated description and the true description. The specific formula is as follows:

[0142]

[0143] where T is the length of the generated description, y t is the word generated at the t-th time step, and finally P(y t |y1,y2…y t-1,F adjusted ) is GPT-4 based on F adjusted The probability of generating the current word from the feature vector and the previous generated content.

[0144] The Adam optimizer was used for training. To prevent instability caused by excessively large learning rates or slow convergence due to excessively small learning rates, a learning rate decay strategy was adopted. The initial learning rate was set to 0.0001, the batch size was set to 32, and training was performed for 20 cycles. L2 regularization was also added during training, and the regularization coefficient was set to 0.0001.

[0145] By optimizing the description generation loss, the training process will continuously improve the accuracy and fluency of the generated descriptions, ultimately achieving the goal of accurately describing the pile.

[0146] Based on the same inventive concept, an embodiment of the present invention provides an industrial material pile identification system.

[0147] The industrial material pile identification system of the present invention can be loaded into an electronic device. According to the functions to be realized, the industrial material pile identification system includes an acquisition module and a processing module.

[0148] The acquisition module can acquire the material pile data to be identified, including radar point cloud data of the material pile, text description data of the stacking status, and environmental data; the processing module can input the material pile data into the material pile type judgment model, and the material pile type judgment model processes the material pile data according to preset rules and outputs the material pile type recognition result;

[0149] The steps for the stockpile type judgment model to process stockpile data are as follows:

[0150] Perform feature extraction on radar point cloud data to obtain point cloud features;

[0151] Perform feature extraction on the stacking status text description data to obtain semantic features;

[0152] Fuse point cloud features and semantic features to obtain fused features;

[0153] Analyze environmental data to determine environmental impact weights, and adjust fusion features based on the environmental impact weights to obtain environmental perception fusion features;

[0154] The pile type recognition result is determined based on the environmental perception fusion features.

[0155] The module described in the present invention may also be referred to as a unit, which refers to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and is stored in a memory of the electronic device.

[0156] The various variations and specific examples of the industrial material pile identification method provided in the above embodiment are also applicable to the industrial material pile identification system of this embodiment. Through the above detailed description of the industrial material pile identification method, those skilled in the art can clearly understand the implementation method of the industrial material pile identification system in this embodiment. For the sake of brevity of the specification, it will not be described in detail here.

[0157] This application also discloses an electronic device, such as Figure 4 Figure 1 is a schematic diagram of the structure of an electronic device for an industrial stockpile identification method according to one embodiment of the present invention. The electronic device may include at least one processor 10, a memory 11 communicatively coupled to the at least one processor, a communication bus 12, and a communication interface 13. It may also include a computer program stored in the memory 11 and executable on the processor 10, such as a method program for industrial stockpile identification.

[0158] In some embodiments, the processor 10 may be composed of an integrated circuit, such as a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting the various components of the entire electronic device using various interfaces and circuits, and executing or executing programs or modules stored in the memory 11 (such as a method for executing industrial stockpile identification) and calling data stored in the memory 11 to perform various functions of the electronic device and process data.

[0159] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 11 can also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Furthermore, the memory 11 can also include both an internal storage unit of the electronic device and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device, such as the code of the method program for industrial material pile identification, but can also be used to temporarily store data that has been output or is to be output.

[0160] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection and communication between the memory 11, the at least one processor 10, etc.

[0161] The communication interface 13 is configured to enable communication between the electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (e.g., a WI-FI interface, a Bluetooth interface, etc.), and is typically configured to establish a communication connection between the electronic device and other electronic devices. The user interface can be a display, an input unit (e.g., a keyboard), and optionally, the user interface can also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch screen, etc. The display can also be referred to as a display screen or a display unit, and is configured to display information processed in the electronic device and to display a visualized user interface.

[0162] Figure 4 Only the electronic device with components is shown, and those skilled in the art can understand that, Figure 4 The structure shown does not limit the electronic device, and the electronic device can include fewer or more components than shown, or combine certain components, or have different component arrangements. For example, although not shown, the electronic device can also include a power supply (e.g., a battery) to supply power to each component. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, so that the power management device can implement functions such as charge management, discharge management, and power consumption management. The power supply can also include one or more direct current or alternating current power supplies, a recharging device, a power supply fault detection circuit, a power supply converter or inverter, a power supply status indicator, etc. The electronic device can also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which are not described here.

[0163] It should be understood that the embodiments are for illustration only, and the scope of the patent application is not limited by the structure.

[0164] Furthermore, if the module / unit integrated into the electronic device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile.

[0165] The present application provides a computer-readable storage medium, including, for example, any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM). The computer-readable storage medium stores a computer program capable of being loaded by a processor and executing the industrial stockpile identification method of the above-described embodiment.

[0166] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "example," "specific example," "one implementation," "a preferred implementation," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0167] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A method for identifying industrial stockpiles, characterized in that: The method comprises: Obtaining data of the material pile to be identified, the material pile data including radar point cloud data of the material pile, text description data of the stacking status, and environmental data; The pile data is input into the pile type judgment model, and the pile type judgment model processes the pile data according to preset rules and outputs the pile type recognition result; The steps for the stockpile type judgment model to process the stockpile data are as follows: Perform feature extraction on radar point cloud data to obtain point cloud features; Perform feature extraction on the stacking status text description data to obtain semantic features; Fuse point cloud features and semantic features to obtain fused features; Analyze environmental data to determine environmental impact weights, and adjust fusion features based on the environmental impact weights to obtain environmental perception fusion features; The pile type recognition result is determined based on the environmental perception fusion features.

2. The industrial stockpile identification method according to claim 1, characterized in that: The feature extraction of the radar point cloud data to obtain point cloud features includes: Input the point cloud data into the multi-layer Transformer encoder to obtain the geometric feature representation of the point cloud data; Multi-scale convolutional networks are used to process geometric feature representation to obtain local and global geometric features of the stockpile. The point cloud features of the stockpile are determined based on local geometric features and global geometric features.

3. The industrial stockpile identification method according to claim 1, characterized in that: The environmental data includes at least one of temperature data, humidity data, and vibration data, and the environmental impact weight is determined according to at least one of the temperature impact factor, the humidity impact factor, and the vibration impact factor; Determine the temperature impact factor of the environment on the pile type based on the temperature data; Determine the humidity impact factor of the environment on the pile shape based on humidity data; Determine the vibration impact factor of the environment on the pile shape based on the vibration data.

4. The industrial stockpile identification method according to claim 3, characterized in that: The calculation formula of temperature influence factor is: in, represents the temperature influence factor; k temp is the temperature sensitivity coefficient, which is used to express the sensitivity of temperature change to the environmental impact weight; E temp Indicates the current temperature value, T0 is the reference temperature value; The calculation formula of humidity impact factor is: in, Represents humidity impact factor; k humidity is the sensitivity coefficient of humidity, which is used to determine the influence of humidity on the environmental impact weight; E humidity is the current humidity value, H0 is the reference humidity value; max(·) means taking the maximum value operation; The calculation formula of vibration influence factor is: in, represents the vibration influence factor; k vibration is the vibration sensitivity coefficient, E vibration is the current vibration amplitude, V0 is the reference vibration value, V threshold is the vibration threshold.

5. The industrial stockpile identification method according to claim 4, characterized in that: The calculation formula for environmental impact weight is: Among them, EnvWeight(E) represents the environmental impact weight, w temp Indicates the weight of the temperature influence factor, w vibration Represents the weight of humidity impact factor, w humidit Indicates the weight of the vibration impact factor.

6. The industrial stockpile identification method according to any one of claims 1 to 5, characterized in that: The method further includes: acquiring a training data set, and using the training data set to train a stockpile type determination model.

7. The industrial stockpile identification method according to claim 6, characterized in that: The training data set includes the pile sample data and the real information of the pile type corresponding to the pile sample data. The training steps of the pile type judgment model include: Construct the network structure of the initial stockpile type judgment model; The network of the initial material pile shape judgment model is trained using the training data set. In each training session, the material pile sample data is input into the initial material pile shape judgment model, and the prediction result of the material pile shape of the material pile sample data by the initial material pile shape judgment model is obtained. The negative log-likelihood loss between the real information and the predicted results is calculated, and the network parameters of the initial material pile shape judgment model are optimized according to the negative log-likelihood loss to obtain the final material pile shape judgment model.

8. An industrial stockpile identification system, used to implement the industrial stockpile identification method according to any one of claims 1 to 7, characterized in that: include: An acquisition module is used to acquire the material pile data to be identified, the material pile data including radar point cloud data of the material pile, text description data of the stacking status and environmental data; A processing module is used to input the stockpile data into the stockpile type judgment model, and the stockpile type judgment model processes the stockpile data according to preset rules and outputs the stockpile type recognition result; The steps for the stockpile type judgment model to process the stockpile data are as follows: Perform feature extraction on radar point cloud data to obtain point cloud features; Perform feature extraction on the stacking status text description data to obtain semantic features; Fuse point cloud features and semantic features to obtain fused features; Analyze environmental data to determine environmental impact weights, and adjust fusion features based on the environmental impact weights to obtain environmental perception fusion features; The pile type recognition result is determined based on the environmental perception fusion features.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor (10); and, a memory (11) communicatively coupled to the at least one processor (10); The memory (11) stores a computer program executable by the at least one processor (10), and the computer program is executed by the at least one processor (10) so that the at least one processor (10) can execute the industrial stockpile identification method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the industrial material pile identification method according to any one of claims 1 to 7 is implemented.