Method and device for determining operation state of robot, storage medium and electronic equipment

By obtaining robot status signal data and using multiple state perception models to classify the robot's operating status, the number of occurrences is counted to determine the robot's operating status, and the problem of inaccurate determination of robot's operating status in the prior art is solved, and the timely and accurate determination of robot's operating status and improvement of work efficiency is achieved.

CN120382486APending Publication Date: 2025-07-29CHINA MOBILE GRP HENAN CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510508067.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art lacks an effective robot operation status determination scheme, and it is impossible to determine the robot operation status in a timely and accurate manner.

Method used

By obtaining the status signal data of the robot, using multiple state perception models to classify the job status, count the number of occurrences of the job status categories output by each state perception model, and determine the job status of the robot based on the categories whose occurrences are greater than or equal to the preset threshold.

Benefits of technology

It realizes the timely and accurate determination of the robot's operating status, and improves the work efficiency and adaptability of the robot in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120382486A_ABST
    Figure CN120382486A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for determining the working state of a robot, a storage medium and electronic equipment, and relates to the technical field of artificial intelligence. According to the state signal data, classifying the operation states of the robot through a plurality of state sensing models, wherein the plurality of state sensing models are obtained by training based on different state signal data of the robot; then determining operation state categories output by all the state sensing models, and counting the occurrence frequency of all the operation state categories; and determining the operation state of the robot according to the operation state category of which the occurrence frequency is greater than or equal to a preset frequency threshold value. According to the technical scheme, the technical problem that an effective robot operation state determining scheme is lacked at present is solved, and the robot operation state can be determined timely and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, device, storage medium, and electronic device for determining the operation state of a robot. Background Art

[0002] With the new generation of information technology revolution, the robot industry plan clearly proposes to promote the integrated application of robots with new technologies such as artificial intelligence, 5G, big data, and cloud computing, improve the intelligence and networking level of robots, and strengthen functional safety, network security, and data security.

[0003] In order to better control the robot, it is necessary to monitor the operation state of the robot. However, there is currently a lack of an effective robot operation state determination scheme, and the operation state of the robot cannot be determined in a timely and accurate manner. Summary of the Invention

[0004] In view of this, the present application provides a method, device, storage medium, and electronic device for determining the operation state of a robot, which solves the technical problem of the lack of an effective robot operation state determination scheme, and can determine the operation state of the robot in a timely and accurate manner.

[0005] In a first aspect, the present application provides a method for determining the operation state of a robot, including:

[0006] Obtaining state signal data of the robot;

[0007] Classifying the operation state of the robot through a plurality of state perception models according to the state signal data, where the plurality of state perception models are respectively trained based on different state signal data of the robot;

[0008] Determining the operation state categories output by each state perception model, and counting the occurrence times of each operation state category;

[0009] Determining the operation state of the robot according to the operation state category whose occurrence times are greater than or equal to a preset number threshold.

[0010] Optionally, determining the operation state categories output by each state perception model, and counting the occurrence times of each operation state category, includes:

[0011] Determining the operation state category with the highest probability output by each state perception model, and counting the occurrence times of the operation state category with the highest probability.

[0012] Optionally, in the case where the occurrence times are all less than the preset number threshold, the method further includes:

[0013] Based on the probabilities of each job status category output by each state perception model, obtain the output result probability distribution corresponding to each state perception model;

[0014] Perform weighted sum analysis according to the output result probability distribution corresponding to each state perception model, and determine the job status of the robot based on the result of the weighted sum analysis.

[0015] Optionally, performing weighted sum analysis according to the output result probability distribution corresponding to each state perception model, and determining the job status of the robot based on the result of the weighted sum analysis includes:

[0016] For any target job status category, obtain the target probabilities output by each state perception model for the target job status category;

[0017] According to the weights corresponding to each state perception model, perform weighted sum on the target probabilities corresponding to each state perception model to obtain the fusion probability of the target job status category;

[0018] Based on the fusion probabilities of each job status category, select the job status category with the largest fusion probability, and determine the job status of the robot based on the job status category with the largest fusion probability.

[0019] Optionally, the method further includes:

[0020] Based on the probabilities of each job status category obtained by visually recognizing the job status of the robot, determine the target probability distribution;

[0021] Calculate the KL divergence differences between the target probability distribution and the output result probability distributions corresponding to each state perception model respectively;

[0022] Determine the weights corresponding to each state perception model according to the KL divergence differences.

[0023] Optionally, the classifying the job status of the robot by using multiple state perception models according to the state signal data includes:

[0024] Obtain a plurality of body signals from the state signal data, and the plurality of body signals include at least one or more of joint pose signals, joint current signals, and joint speed signals;

[0025] Use an encoder based on Transformer to convert the plurality of body signals into corresponding embedding vectors respectively;

[0026] Input the embedding vectors of the plurality of body signals into a state perception model based on the fusion of the plurality of body signals to obtain the probabilities of each job status category; and,

[0027] The embedding vector of each ontology signal is input into the corresponding state perception model based on a single ontology signal to obtain the probability of each job status category.

[0028] Optionally, the training process of the state perception model based on the fusion of multiple ontology signals includes:

[0029] Acquire multiple sample body signals of the robot under different working tasks, wherein the multiple sample body signals include at least one or more of sample joint posture signals, sample joint current signals, and sample joint velocity signals;

[0030] constructing positive sample pairs based on the multiple sample ontology signals in the same operating state, and constructing negative sample pairs based on the multiple sample ontology signals in different operating states;

[0031] Based on the positive sample pairs and the negative sample pairs, the state perception model based on the fusion of multiple ontological signals is trained by minimizing the contrast loss function to learn to bring the embedding vectors of the positive sample pairs closer and separate the embedding vectors of the negative sample pairs.

[0032] Optionally, the training process of the state perception model based on a single ontology signal includes:

[0033] Acquire a single sample body signal of the robot under different working tasks, wherein the single sample body signal is one of a sample joint posture signal, a sample joint current signal, and a sample joint velocity signal;

[0034] constructing positive sample pairs based on the single sample ontology signals at different times under the same operating state, and constructing negative sample pairs based on the single sample ontology signals under different operating states;

[0035] Based on the positive sample pairs and the negative sample pairs, the state perception model based on the single ontology signal is trained by minimizing the contrast loss function to learn to bring the embedding vectors of the positive sample pairs closer and separate the embedding vectors of the negative sample pairs.

[0036] In a second aspect, the present application provides a device for determining the operating status of a robot, comprising:

[0037] an acquisition module, configured to acquire status signal data of the robot;

[0038] a classification module configured to classify the robot's operating status according to the status signal data using a plurality of status perception models, wherein the plurality of status perception models are trained based on different status signal data of the robot;

[0039] A determination module, configured to determine the job status categories output by each status perception model, and count the occurrence times of each of the job status categories; and determine the job status of the robot according to the job status categories whose occurrence times are greater than or equal to a preset number threshold.

[0040] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0041] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the method described in the first aspect is implemented.

[0042] In a fifth aspect, the present application provides a computer program product, on which a computer program is stored, and when the computer program product is executed by a processor, the method described in the first aspect is implemented.

[0043] By means of the above technical solutions, the present application provides a method, device, storage medium, and electronic device for determining the job status of a robot. The method includes: first, obtaining the status signal data of the robot; then, classifying the job status of the robot through multiple status perception models according to the status signal data, and the multiple status perception models are respectively trained based on different status signal data of the robot; then, determining the job status categories output by each status perception model, and counting the occurrence times of each job status category; and further determining the job status of the robot according to the job status categories whose occurrence times are greater than or equal to a preset number threshold. By applying the technical solutions of the present application, the technical problem of the lack of an effective solution for determining the job status of a robot at present is solved, and the job status of the robot can be determined in a timely and accurate manner.

[0044] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically described below. Description of the Drawings

[0045] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 The flowchart shows a method for determining the operation state of a robot provided by an embodiment of the present application;

[0048] Figure 2 The flowchart shows another method for determining the operation state of a robot provided by an embodiment of the present application;

[0049] Figure 3 The flowchart shows an example flowchart provided by an embodiment of the present application;

[0050] Figure 4 The structural diagram shows a device for determining the operation state of a robot provided by an embodiment of the present application. Detailed implementation manners

[0051] The embodiments of the present application will be described in more detail with reference to the accompanying drawings. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0052] To solve the technical problem of the current lack of an effective solution for determining the operation state of a robot, an embodiment of the present application provides a method for determining the operation state of a robot, as Figure 1 shown, the method includes:

[0053] Step S101, obtain the state signal data of the robot.

[0054] When the robot performs different types of operations, the state signal data of its body will change in real time, and these changes in the state signal data can be used as the classification basis for the operation state of the robot.

[0055] Among them, the state signal data of the robot may include at least one of the robot joint pose signal, the robot joint current signal, and the robot joint speed signal, etc.

[0056] For the execution subject of this embodiment, it can be a device or equipment for determining the operation state of the robot, and the operation state of the robot can be accurately determined (or can be called perceived, etc.) based on the state signal data of the robot body.

[0057] Step S102, classify the operation state of the robot through multiple state perception models according to the state signal data of the robot.

[0058] Multiple state perception models are separately trained based on different state signal data of the robot. Each state perception model among these state perception models can be used to classify the working state of the robot to obtain the classification result of the robot's working state.

[0059] Embodiments of the present disclosure can calculate through multiple state perception models according to the state signal data of the robot, and respectively obtain the working state of the robot corresponding to each state perception model.

[0060] Step S103: Determine the working state categories output by each state perception model, and count the occurrence times of each working state category.

[0061] For example, each state perception model can output the working state category it predicts. For example, state perception model A among the multiple state perception models outputs its predicted working state category a, state perception model B outputs its predicted working state category b, state perception model C outputs its predicted working state category a, state perception model D outputs its predicted working state category a, and state perception model E outputs its predicted working state category a. Then the occurrence times of the working state category a is 4 times, and the occurrence times of the working state category b is 1 time.

[0062] Step S104: Determine the working state of the robot according to the working state category whose occurrence times are greater than or equal to the preset times threshold.

[0063] The preset times threshold can be determined according to the number of state perception models, and this preset times threshold can be used to determine the majority of working state categories. For example, if there are 4 state perception models, the preset times threshold can be 3. That is, after the statistics in step S103, if there is a working state category whose occurrence times are greater than or equal to 3 times, then determine the working state of the robot according to this working state category.

[0064] By applying the technical solution of the embodiments of the present application, the technical problem of the lack of an effective robot working state determination solution at present is solved, and the working state of the robot can be determined in a timely and accurate manner.

[0065] To further illustrate the implementation process of the above embodiments, a specific implementation manner is given, as Figure 2 shown, including:

[0066] Step S201: Obtain the state signal data of the robot.

[0067] Among them, the state signal data of the robot may include at least one of the robot joint pose signal, the robot joint current signal, and the robot joint speed signal, etc.

[0068] Step S202: Classify the robot's operating status respectively through multiple state perception models according to the status signal data.

[0069] The multiple state perception models can be trained separately based on different state signal data of the robot.

[0070] In some embodiments, the multiple state perception models may include a state perception model based on the fusion of multiple body signals of the robot, and a state perception model based on a single body signal of the robot, wherein each body signal may have its own corresponding state perception model. Accordingly, step S202 may specifically include: first, obtaining multiple body signals from the state signal data of the robot, wherein the multiple body signals include at least one or more of the joint posture signal, the joint current signal, and the joint velocity signal; then using a Transformer-based encoder to convert the multiple body signals into corresponding embedding vectors; then inputting the embedding vectors of the multiple body signals into the state perception model based on the fusion of multiple body signals to obtain the probability of each work state category; and, inputting the embedding vector of each body signal into the state perception model based on the corresponding single body signal to obtain the probability of each work state category.

[0071] In some examples, the training process of a state perception model based on the fusion of multiple ontological signals includes: first obtaining multiple sample ontological signals of the robot under different working tasks, and the multiple sample ontological signals include at least one or more of sample joint posture signals, sample joint current signals and sample joint velocity signals; then constructing positive sample pairs based on multiple sample ontological signals under the same working state, and constructing negative sample pairs based on multiple sample ontological signals under different working states; then, based on the constructed positive sample pairs and negative sample pairs, the state perception model based on the fusion of multiple ontological signals is trained by minimizing the contrast loss function to learn to bring the embedding vectors of the positive sample pairs closer and separate the embedding vectors of the negative sample pairs.

[0072] For example, when a robot performs different types of operations, its state signal data will change in real time. These changes in state signals can be used as a basis for classifying the robot's operating status. Figure 3 As shown in the figure, firstly, the body signals of the multifunctional robot under different working tasks are collected, including the robot joint posture signal, robot joint current signal, and robot joint speed signal, which are represented as S and p (n p / T p ), S c (n c / T c ), S v (n v / Tv ), where T is the sampling frequency and n is the sampling point number.

[0073] S p (n p / T p ), S c (n c / T c ), S v (n v / T v ) are all different time series signals during the robot's operation. Therefore, any one of the signals can characterize the robot's operation state at that time series in the time series.

[0074] To make full use of different types of robot body signals, the embodiments of the present application can construct a state perception model from multiple perspectives of body signal fusion, which is specifically as follows:

[0075] The embodiments of the present application can construct a state classification model for multi-body signal contrast learning, that is, a state perception model based on the fusion of multiple body signals. For the convenience of representation, hereinafter, S p (n p / T p ), S c (n c / T c ), S v (n v / T v ) will be uniformly represented as S p , S c , S v .

[0076] The first step is to define the input space:

[0077] Each signal is composed of features with d input dimensions within T time steps. Use the Transformer-based encoder network f θ to map these signals into feature representations:

[0078] X i = f θ (S i ), i = p, c, v

[0079] where S i ∈R T ×d input represents a signal sequence with T time steps, and each time step has d input features.

[0080] Next, construct the Transformer encoder network structure for the robot operation time series signal:

[0081] (1) Input layer: Each signal S p 、S c 、S v has an input dimension of T×d input , that is, it contains T time steps, and the feature dimension of each time step is d input .

[0082] (2) Positional encoding: Since the Transformer model has no inherent ability to perceive the time series order, positional encoding PE needs to be added to the input signal. The expression of positional encoding is as follows:

[0083]

[0084] Then, the positional encoding is added element-wise to the signal S i to obtain the signal S i pos with positional information:

[0085]

[0086] (3) Multi-head self-attention mechanism: The global dependencies between time steps of the signal are modeled through the multi-head self-attention mechanism. The calculation process of self-attention is as follows:

[0087]

[0088] where Q, K, and V are the query, key, and value vectors respectively, which are obtained by linearly transforming the input signal. Multi-head self-attention calculates the results of multiple attention heads in parallel and finally concatenates them:

[0089] MultiHead(Q,K,V)=Concat(head1,…,head h )W O

[0090] (4) Feed-forward network: After each self-attention layer, two fully connected layers and the non-linear activation function ReLU are applied. The specific formula is as follows:

[0091] FFN(x)=ReLU(xW1+b1)W2+b2

[0092] (5) Residual connection and layer normalization: Residual connection and layer normalization are used to enhance the training stability of the model:

[0093] LayerNorm(x+Attention(x)),LayerNorm(x+FFN(x))

[0094] (6) Output layer: After passing through multiple Transformer encoder layers, the obtained embedding representation where dmodel is the output embedding dimension of the model. To obtain global features, the signal can be aggregated in the time dimension through global average pooling:

[0095]

[0096] At this time, the signal S p 、S c 、S v are all transformed into corresponding embedding vectors X p 、X c 、X v .

[0097] Second step, construct positive and negative sample pairs:

[0098] In the contrastive learning of multiple robot body signals, the construction of sample pairs is very crucial. The construction logic of sample pairs is as follows:

[0099] In the same state: The signals S p 、S c 、S v collected within the same time step should be regarded as different dimensional information in the same state. Therefore, their embedding vectors should be close in the feature space.

[0100] In different states: The signals collected at different time steps, especially the information from different operation states, should be regarded as negative samples. Even if the same type of signal (such as S p ) is collected in different operation states, their embedding vectors should also be separated in the feature space.

[0101] Therefore, the positive and negative sample pairs are defined as follows:

[0102] (1) Positive sample pairs: The construction of positive sample pairs is based on the collection of multiple signals within the same time step or the same period of time. At a certain time step t, if the signals S p 、S c 、S v are signals collected simultaneously from the robot in the same operation state, then it is considered that these signals should be close in the embedding space. The positive sample is defined as:

[0103]

[0104] where are the signal embedding vectors from the same time step t. At this time, it can be assumed that these signals are correlated, so their embeddings should be close in the contrastive learning.

[0105] (2) Negative sample pairs: Negative sample pairs come from signals in different operating states. During the same data acquisition process, different time steps may correspond to different operating states. For example, at time step t1, the robot is in operating state A, and at time step t2, the robot is in operating state B. At this time, the signal from state A and the signal from state B should be distinguished in the embedding space. Negative sample pairs can be:

[0106]

[0107] These negative sample pairs are signal embedding vectors from different time steps (i.e., different job states), so they should be far away from each other in the embedding space.

[0108] By constructing positive and negative sample pairs in this way, the model can better capture the signal differences of the robot in different working states through comparative learning, and ultimately improve the accuracy of state classification.

[0109] The third step is to define the contrast loss function:

[0110] A contrastive loss function based on InfoNCE is used. Its goal is to learn useful signal embeddings by minimizing the distance between positive sample pairs and maximizing the distance between negative sample pairs. The expression of the contrastive loss function is:

[0111]

[0112] Where: sin(X i ,X j ) is the cosine similarity, which is used to measure the similarity between embedding vectors, and τ is the temperature parameter that controls the smoothness of the loss distribution:

[0113]

[0114] Step 4: Training the model:

[0115] By minimizing the contrastive loss function L contrastive , train the Transformer-based encoder network f θ During training, the model learns to bring the embedding vectors of positive pairs closer together and the embedding vectors of negative pairs farther apart. This allows the model to capture both similarities and differences between different signals.

[0116] Step 5: Status classification:

[0117] After the above contrastive learning network training is completed, the signal S p , S c , S v The embedding representation X p , X c , Xv It can be used for state classification tasks. The network for the classification task is a simple MLP network. This lightweight classification network is trained using cross entropy loss to map these embedded vectors to specific state labels yi.

[0118] Through the above process, the global dependency between different signals of robot operations can be effectively learned, and the state classification model can be optimized through comparative learning.

[0119] In some examples, the training process of a state perception model based on a single ontology signal includes: first obtaining a single sample ontology signal of the robot under different working tasks, where the single sample ontology signal is one of a sample joint posture signal, a sample joint current signal, and a sample joint velocity signal; then constructing a positive sample pair based on the single sample ontology signal at different times under the same working state, and constructing a negative sample pair based on the single sample ontology signal under different working states; then, based on the positive sample pair and the negative sample pair, the state perception model based on the single ontology signal is trained by minimizing the contrast loss function to learn to bring the embedding vectors of the positive sample pair closer and separate the embedding vectors of the negative sample pair.

[0120] For example, in order to fully utilize different types of robot proprioceptive signals, the embodiment of the present disclosure can construct a state perception model from the perspective of a single proprioceptive signal, as described below:

[0121] During a robot's operation, a single signal can directly reflect the robot's operating status. To fully utilize the characteristic differences of time series signals in the time dimension, positive and negative samples are constructed for each signal. This process is based on the changes in the signal at different time steps or states of the robot. The specific construction process is as follows:

[0122] Take the robot joint posture signal Sp as an example to illustrate. The process of the robot joint current signal Sc and the robot joint velocity signal Sv is the same:

[0123] (1) Positive sample pairs: For a single signal type, construct positive sample pairs based on different time steps under the same state. and Both come from the same working state (for example, the robot works in the same working mode), then the signals of these two time steps can constitute a positive sample pair:

[0124]

[0125] In this way, signals from the same job state are embedded in the vector and should be close in the embedding space.

[0126] (2) Negative sample pairs: Negative sample pairs are constructed based on the signal time steps in different states. From job state A, and the signal From job state B, the signals of these two time steps should constitute a negative sample pair. The definition of a negative sample pair is:

[0127]

[0128] This means that the signals collected under different working conditions are embedded in the vector and should be separated in the embedding space.

[0129] (3) Model training and state classification: After constructing the positive and negative sample pairs, the contrastive loss function is used to optimize the model so that the signals from the same job state are closer and the signals from different job states are more separated. The loss function is as follows:

[0130]

[0131] in: and is a positive sample pair from the same state; and are all possible signal pairs, including negative sample pairs; τ is the temperature parameter; is the cosine similarity, which is used to measure the similarity between two signal embeddings.

[0132] Finally, by minimizing the contrastive loss function, the model learns how to bring the embedding vectors of signals in the same operating state closer together and separate the embedding vectors of signals in different operating states. The network for the classification task is a simple MLP network. This lightweight classification network is trained using cross-entropy loss to map these embedding vectors to specific state labels yi.

[0133] Through contrastive learning, the model is able to learn the changes and differences in job status from a single signal type, thereby performing better in classification tasks.

[0134] Step S203: Determine the job status category with the maximum probability output by each state perception model, and count the number of occurrences of the job status category with the maximum probability.

[0135] Each state perception model can output the probabilities of the robot for each job state category. For example, among multiple state perception models, state perception model 1 can output the probabilities of the robot for each job state category, and state perception model 2 can also output the probabilities of the robot for each job state category. For the probabilities of each job state category output by state perception model 1, select the job state category with the maximum probability; similarly, for the probabilities of each job state category output by state perception model 2, select the job state category with the maximum probability. Based on this method, count the occurrence times of the job state categories with the maximum probabilities output by each state perception model. For example, if the job state categories with the maximum probabilities output by 8 state perception models are all job state a, then the occurrence times corresponding to job state a is 8 times; if the job state categories with the maximum probabilities output by 6 state perception models are all job state b, then the occurrence times corresponding to job state b is 6 times.

[0136] Step S204a: If there exists a job state category whose occurrence times are greater than or equal to the preset times threshold, then determine the job state of the robot according to the job state category whose occurrence times are greater than or equal to the preset times threshold.

[0137] For example, as Figure 3 shown, assume that there are L categories of the robot's job states, denoted as C1, C2, …, C L For each category C l (l = 1, 2, ..., L), according to the construction of the above 4 state perception classification models, the state perception classification results of the robot are as follows:

[0138] ① The probability results of the state recognition of the fused signal are P m (C1), P m (C2), …, P m (C L ).

[0139] ② The probability results of the state recognition based on the joint pose are P p (C1), P p (C2), …, P p (C L ).

[0140] ③ The probability results of the state recognition based on the current state are P c (C1), P c (C2), …, P c (C L ).

[0141] ④ The probability results of the state recognition based on the joint velocity are P v (C1), P v (C2), …, P v(C L )。

[0142] For the above four state perception models, select the category by the maximum probability principle:

[0143]

[0144] In this way, each of the four models selects a category as the final classification result according to their output probabilities.

[0145] Define a counter vector count, where each element count[l] represents the number of times the category C l appears in the outputs of the four models.

[0146]

[0147] where, is an indicator function. If the output category of model i is Cl, its value is 1; otherwise, it is 0.

[0148] If the count of a certain category C l count[l] ≥ 3, directly output this category as the final category.

[0149] If count[l] ≥ 3

[0150] Otherwise, it is necessary to recalculate the category probability using visual assistance, that is, perform the processes shown in steps S204b to S205b.

[0151] Step S204b parallel to step S204a, in the case where the occurrence times are all less than the preset number threshold, based on the probabilities of each job state category output by each state perception model, obtain the output result probability distribution corresponding to each state perception model.

[0152] Step S205b, perform weighted sum analysis according to the output result probability distribution corresponding to each state perception model, and determine the operation state of the robot according to the result of the weighted sum analysis.

[0153] In some embodiments, step S205b may specifically include: for any target job state category, obtain the target probabilities output by each state perception model for the target job state category respectively; then, according to the weights corresponding to each state perception model, perform weighted sum on the target probabilities corresponding to each state perception model to obtain the fusion probability of the target job state category; then, based on the fusion probabilities of each job state category, select the job state category with the largest fusion probability, and determine the operation state of the robot according to the job state category with the largest fusion probability.

[0154] For example, as Figure 3 shown, for the probability distribution of the L category for the results output by the above 4 models in the correction of robot state assessment based on visual assistance, that is, for category C l , each model will give the corresponding probability. Then for category C l , the probability results of state recognition of the above 4 state perception classification models are respectively: P m (C l ), P p (C l ), P c (C l ), P v (C l ).

[0155] Weighted voting mechanism: To fuse these results, a weighted voting mechanism is adopted to sum the 4 results with weights to obtain the final category probability distribution. Let the weights of each model be w m , w p , w c , w v , then the fusion probability P(C l ) of category C l is:

[0156] P(C l ) = w m P m (C l ) + w p P p (C l ) + w c P c (C l ) + w v P v (C l )

[0157] After obtaining the fusion probabilities P(C l ), P(C2),..., P(C L ), select the category with the maximum probability as the final recognition result. That is, the final category is:

[0158]

[0159] Through this process, the fused probability values can be converted into specific categories.

[0160] Further optionally, a target probability distribution may be determined based on the probabilities of each job status category obtained by visually recognizing the robot's job status; then, the KL divergence differences between the target probability distribution and the output result probability distributions corresponding to each state perception model are calculated; and then, the weights corresponding to each state perception model are determined based on the KL divergence differences.

[0161] For example, assume that the result of visual signal recognition is P 视觉 (C l ), that is, the probability distribution of each category given by the visual signal.

[0162] To update the model weights based on the visual results, first calculate the difference between the visual signal and the output of each model. The difference between the probability distribution P m (C l ), P p (C l ), P c (C l ), P v (C l ) of each model and the visual result can be calculated by a certain distance metric (such as KL divergence).

[0163] The KL divergence is defined as:

[0164] For each model, calculate the KL divergence difference between its probability distribution and the visual signal:

[0165] ① The difference between the combined signal and the visual signal: D KL (P 视觉 ||P m (C l ))

[0166] ② The difference between the joint pose and the visual signal: D KL (P 视觉 ||P p (C l ))

[0167] ③ The difference between the flow and the visual signal: D KL (P 视觉 ||P c (C l ))

[0168] ④ The difference between the joint speed and the visual signal: D KL (P 视觉 ||P v (C l ))

[0169] Based on these differences, update the weights, and the update rule is:

[0170] Then, normalize the updated weights:

[0171] In this way, gradually increase the weights of the models that are closer to the visual signal results, making the above weighted voting mechanism more biased towards the reliable model outputs. That is, compare the visual signal with the outputs of each state perception model through KL divergence and adjust the weights of the models. After the weights are updated, the models with higher weights will have a greater impact on the final decision. This is equivalent to "weighting" the outputs of each model through the visual signal and enhancing the influence of the models that are closer to the visual signal results.

[0172] In some examples, the job state perception of the robot mainly depends on the external sensors installed on the robot body, but it is difficult to comprehensively and real-time obtain the state information of the robot during the switching process of different processing functions. Especially in a complex dynamic environment, the external sensors may not be able to respond in time or provide sufficient information support. This limitation affects the working efficiency and adaptability of the multi-functional processing robot. And the robot perception technology represented by vision technology has characteristics such as large amount of data and slow calculation speed, which is difficult to meet the real-time requirement of continuous state monitoring.

[0173] For this reason, the embodiments of the present application provide a method for multi-functional robot job state perception based on ontology signals and vision correction. That is, as Figures 1 to 3 shown in the embodiment content, use the pose-current-velocity signals to form sample pairs by using the dual factors of time and state to construct a perception model, evaluate the robot state based on the voting mechanism, and combine vision assistance to perform state evaluation and correction. By integrating the multi-source information of the robot ontology and supplemented by vision information, reduce the deployment difficulty and cost of sensors, and solve the problem of resource consumption in traditional training using vision information, and realize the full-process perception of the job process of the multi-functional processing robot. By applying the technical solution of the present application, the working efficiency and adaptability of the robot in a complex manufacturing environment can be improved, providing more advanced and reliable technical support and commercial value for intelligent manufacturing.

[0174] Further, as a Figures 1 to 3 specific implementation of the method shown, the embodiments of the present application provide a device for determining the job state of the robot, as Figure 4 shown, the device includes: an acquisition module 301, a classification module 302, and a determination module 303.

[0175] The acquisition module 301 is configured to acquire the state signal data of the robot;

[0176] A classification module 302, configured to classify the operation states of the robot respectively through a plurality of state perception models according to the state signal data, where the plurality of state perception models are respectively trained based on different state signal data of the robot;

[0177] A determination module 303, configured to determine the operation state categories output by each state perception model, and count the occurrence times of each of the operation state categories; and determine the operation state of the robot according to the operation state categories whose occurrence times are greater than or equal to a preset times threshold.

[0178] In some examples of this embodiment, the determination module 303 is specifically configured to determine the operation state category with the highest probability output by each state perception model, and count the occurrence times of the operation state category with the highest probability.

[0179] In some examples of this embodiment, the determination module 303 is further configured to, when the occurrence times are all less than the preset times threshold, obtain the output result probability distribution corresponding to each state perception model based on the probabilities of each operation state category output by each state perception model; perform weighted summation analysis according to the output result probability distribution corresponding to each state perception model, and determine the operation state of the robot according to the weighted summation analysis result.

[0180] In some examples of this embodiment, the determination module 303 is specifically configured to, for any target operation state category, obtain the target probabilities output by each state perception model for the target operation state category respectively; perform weighted summation on the target probabilities corresponding to each state perception model according to the weights corresponding to each state perception model to obtain the fusion probability of the target operation state category; based on the fusion probabilities of each operation state category, select the operation state category with the highest fusion probability, and determine the operation state of the robot according to the operation state category with the highest fusion probability.

[0181] In some examples of this embodiment, the determination module 303 is further configured to determine a target probability distribution based on the probabilities of each operation state category obtained by performing visual signal recognition on the operation state of the robot; calculate the KL divergence difference between the target probability distribution and the output result probability distribution corresponding to each state perception model respectively; and determine the weights corresponding to each state perception model according to the KL divergence difference.

[0182] In some examples of this embodiment, the classification module 302 is specifically configured to obtain multiple ontological signals from the state signal data, wherein the multiple ontological signals include at least one or more of joint posture signals, joint current signals, and joint velocity signals; use a Transformer-based encoder to convert the multiple ontological signals into corresponding embedding vectors; input the embedding vectors of the multiple ontological signals into a state perception model based on the fusion of multiple ontological signals to obtain the probability of each work state category; and input the embedding vector of each ontological signal into each corresponding state perception model based on a single ontological signal to obtain the probability of each work state category.

[0183] In some examples of this embodiment, the device also includes a training module, wherein the training module includes the following steps for the training process of the state perception model based on the fusion of multiple ontological signals: obtaining multiple sample ontological signals of the robot under different working tasks, wherein the multiple sample ontological signals include at least one or more of sample joint posture signals, sample joint current signals and sample joint speed signals; constructing positive sample pairs based on the multiple sample ontological signals under the same working state, and constructing negative sample pairs based on the multiple sample ontological signals under different working states; based on the positive sample pairs and the negative sample pairs, training the state perception model based on the fusion of multiple ontological signals is performed by minimizing the contrast loss function to learn to bring the embedding vectors of the positive sample pairs closer and separate the embedding vectors of the negative sample pairs.

[0184] In some examples of this embodiment, the training module for the training process of the state perception model based on a single ontology signal includes: obtaining a single sample ontology signal of the robot under different working tasks, wherein the single sample ontology signal is one of a sample joint posture signal, a sample joint current signal, and a sample joint speed signal; constructing a positive sample pair based on the single sample ontology signal at different times under the same working state, and constructing a negative sample pair based on the single sample ontology signal under different working states; based on the positive sample pair and the negative sample pair, training the state perception model based on a single ontology signal by minimizing the contrast loss function to learn to bring the embedding vectors of the positive sample pair closer and separate the embedding vectors of the negative sample pair.

[0185] It should be noted that for other corresponding descriptions of the functional units involved in the device for determining the operation status of a robot provided in the embodiment of the present application, please refer to Figures 1 to 3 The corresponding description in will not be repeated here.

[0186] Based on the above Figures 1 to 3The method described above. Correspondingly, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-described method as Figures 1 to 3 shown is implemented.

[0187] Based on the above method as Figures 1 to 3 shown. Correspondingly, this embodiment also provides a computer program product, on which a computer program is stored. When the computer program product is executed by a processor, the above-described method as Figures 1 to 3 shown is implemented.

[0188] Based on such an understanding, the technical solution of the embodiment of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of the present application.

[0189] Based on the above method as Figures 1 to 3 shown, and Figure 4 the virtual device embodiment shown. To achieve the above object, an embodiment of the present application also provides an electronic device, such as a terminal device or a server, etc. The electronic device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the above method as Figures 1 to 3 shown.

[0190] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.

[0191] Those skilled in the art can understand that the above-described physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0192] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between the components inside the storage medium, and communication between other hardware and software in the information processing physical device.

[0193] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. The embodiments of this application provide a multifunctional robot operation state perception method based on ontology signals and visual correction. That is Figures 1 to 3 As shown in the embodiment content, sample pairs are formed by using pose-current-velocity signals with dual factors of time and state to construct a perception model, and the robot state is evaluated based on a voting mechanism, and visual assistance is combined for state evaluation correction. By integrating multi-source information of the robot body and supplementing it with visual information, the deployment difficulty and cost of sensors are reduced, and the problem of resource consumption in traditional visual information training is solved, realizing the full-process perception of the operation process of the multifunctional processing robot. By applying the technical solution of this application, the working efficiency and adaptability of the robot in a complex manufacturing environment can be improved, providing more advanced and reliable technical support and commercial value for intelligent manufacturing.

[0194] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0195] The above are only specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments described herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for determining the working state of a robot, characterized in that Including: Obtain the status signal data of the robot; Classify the working status of the robot respectively through multiple status perception models according to the status signal data, and the multiple status perception models are respectively trained based on different status signal data of the robot; Determine the working status categories output by each status perception model, and count the occurrence times of each working status category; Determine the working status of the robot according to the working status category whose occurrence times are greater than or equal to the preset times threshold.

2. The method according to claim 1, wherein Determine the working status categories output by each status perception model, and count the occurrence times of each working status category, including: Determine the working status category with the highest probability output by each status perception model, and count the occurrence times of the working status category with the highest probability.

3. The method according to claim 1, characterized in that, In the case that the occurrence times are all less than the preset times threshold, the method further includes: Based on the probabilities of each working status category output by each status perception model, obtain the output result probability distribution corresponding to each status perception model; Perform weighted sum analysis according to the output result probability distribution corresponding to each status perception model, and determine the working status of the robot according to the weighted sum analysis result.

4. The method according to claim 3, characterized in that, Perform weighted sum analysis according to the output result probability distribution corresponding to each status perception model, and determine the working status of the robot according to the weighted sum analysis result, including: For any target working status category, obtain the target probabilities output by each status perception model for the target working status category respectively; According to the weights corresponding to each status perception model, perform weighted sum on the target probabilities corresponding to each status perception model to obtain the fusion probability of the target working status category; Based on the fusion probabilities of each working status category, select the working status category with the highest fusion probability, and determine the working status of the robot according to the working status category with the highest fusion probability.

5. The method according to claim 4, wherein The method further includes: Based on the probabilities of each working status category obtained by visual signal recognition of the robot's working status, determine the target probability distribution; Calculate the KL divergence difference between the target probability distribution and the output result probability distribution corresponding to each status perception model respectively; Determine the weights corresponding to each status perception model according to the KL divergence difference.

6. The method according to claim 1, wherein The classifying the working status of the robot respectively through multiple status perception models according to the status signal data includes: Obtain multiple body signals from the status signal data, and the multiple body signals include at least one or more of joint pose signals, joint current signals and joint speed signals; Use an encoder based on Transformer to convert the multiple body signals into corresponding embedding vectors respectively; Input the embedding vectors of the multiple body signals into a status perception model based on the fusion of multiple body signals to obtain the probabilities of each working status category; and Input the embedding vector of each body signal into the corresponding status perception model based on a single body signal respectively to obtain the probabilities of each working status category.

7. The method according to claim 6, wherein The training process of the state perception model based on the fusion of multiple ontological signals includes: Acquire multiple sample body signals of the robot under different working tasks, wherein the multiple sample body signals include at least one or more of sample joint posture signals, sample joint current signals, and sample joint velocity signals; constructing positive sample pairs based on the multiple sample ontology signals in the same operating state, and constructing negative sample pairs based on the multiple sample ontology signals in different operating states; Based on the positive sample pairs and the negative sample pairs, the state perception model based on the fusion of multiple ontological signals is trained by minimizing the contrast loss function to learn to bring the embedding vectors of the positive sample pairs closer and separate the embedding vectors of the negative sample pairs.

8. The method according to claim 6, characterized in that, The training process of the state perception model based on a single ontology signal includes: Acquire a single sample body signal of the robot under different working tasks, wherein the single sample body signal is one of a sample joint posture signal, a sample joint current signal, and a sample joint velocity signal; constructing positive sample pairs based on the single sample ontology signals at different times under the same operating state, and constructing negative sample pairs based on the single sample ontology signals under different operating states; Based on the positive sample pairs and the negative sample pairs, the state perception model based on the single ontology signal is trained by minimizing the contrast loss function to learn to bring the embedding vectors of the positive sample pairs closer and separate the embedding vectors of the negative sample pairs.

9. A device for determining the working state of a robot, characterized in that include: an acquisition module, configured to acquire status signal data of the robot; a classification module configured to classify the robot's operating status according to the status signal data using a plurality of status perception models, wherein the plurality of status perception models are trained based on different status signal data of the robot; a determination module configured to determine the job status category output by each state perception model and count the number of occurrences of each job status category; The operating state of the robot is determined based on the operating state category whose occurrence number is greater than or equal to a preset number threshold.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

11. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

12. A computer program product, on which a computer program is stored, characterized in that, When the computer program product is executed by a processor, the method according to any one of claims 1 to 8 is implemented.