Device state recognition method and apparatus

CN122761027APending Publication Date: 2026-09-15GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610863446.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-15

Smart Images

  • Figure CN122761027A_ABST
    Figure CN122761027A_ABST
Patent Text Reader

Abstract

The application relates to a device state recognition method and device, the method comprising: obtaining each text word element associated with global visual features and global text features of a target device. According to the global visual features, a target word element is selected from the text word elements. According to the target word element, a text global representation is determined. According to the global visual features and the text global representation, a state recognition result of the target device is determined. Through a visual-guided text word element screening mechanism, invalid text information is eliminated, the text semantic structure is refined, the accuracy and efficiency of cross-modal feature alignment are significantly improved, redundant calculation is reduced, the model response speed is accelerated, the consistency of the image scene and the text semantic is strengthened, and the accuracy and robustness of the device state recognition in a complex power distribution network environment are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power equipment technology, and in particular to a method and apparatus for identifying equipment status. Background Technology

[0002] With the accelerated development of smart grids, the distribution network, as a crucial link connecting the transmission network and end users, is receiving increasing attention for its operational reliability, power quality, and safety. Among these, 10kV overhead lines constitute the backbone of the distribution network, widely distributed throughout urban and rural areas, and their operational status directly affects the normal order of social production and daily life.

[0003] Currently, most methods for locating faults and identifying the status of power equipment such as connectors, fuses, insulators, transformers, and overhead switches rely on images or text information fed back from camera equipment, power drone inspections, and autonomous power robot operations. However, current equipment status identification methods based on image or text recognition suffer from significant drawbacks. The large volume of text information not only increases the recognition burden but also results in slow response times and inaccurate identification results. Summary of the Invention

[0004] Therefore, it is necessary to provide a method and apparatus for power equipment status identification that offers both efficiency and accuracy in addressing the aforementioned technical problems.

[0005] Firstly, this application provides a device status identification method. The method includes:

[0006] Obtain the text lexical units associated with the global visual features and global text features of the target device;

[0007] Based on global visual features, target lexical units are selected from each text lexical unit.

[0008] Determine the global representation of the text based on the target lexical units;

[0009] The state recognition result of the target device is determined based on global visual features and global text representation.

[0010] In one embodiment, selecting target lexical units from various text lexical units based on global visual features includes:

[0011] Global average pooling is applied to the global visual features to obtain a global visual representation.

[0012] For each text word, the text word is concatenated with the shown visual global representation to obtain the concatenated feature;

[0013] The target score for each text word is determined based on the splicing features and visual global representation corresponding to each text word.

[0014] Based on the target score of each text word, target words are selected from each text word.

[0015] In one embodiment, the target score for each text word is determined based on the concatenation features and visual global representation corresponding to each text word, including:

[0016] For each text word, a basic score for the text word is obtained based on the multilayer perceptron and the concatenation features corresponding to the text word.

[0017] Based on the text prior mapping function and text words, determine the prior score of the text words;

[0018] The target score for each word in the text is determined based on the base score and prior score.

[0019] In one embodiment, selecting target words from each text word based on the target score of each text word includes:

[0020] Based on the target score, the text words are sorted to obtain the sorting results;

[0021] Based on the sorting results, target terms are selected from each text terminology.

[0022] In one embodiment, determining the global representation of the text based on the target lexical units includes:

[0023] Based on the target score of each target word, determine the word weight of each target word;

[0024] Based on the lexical weights of each target lexical, a weighted sum is calculated for each target lexical to obtain a global representation of the text.

[0025] In one embodiment, the state recognition result of the target device is determined based on global visual features and global text representation, including:

[0026] Based on global visual features, global text features, and global text representation, determine cross-modal fusion features;

[0027] Based on cross-modal fusion features and global visual features, residual fusion features are determined;

[0028] Based on the residual fusion characteristics, the state recognition result of the target device is determined.

[0029] In one embodiment, cross-modal fusion features are determined based on global visual features, global text features, and global text representation, including:

[0030] Based on global visual features, determine the sequential visual features;

[0031] Attention weights are determined based on sequential visual features and global text features;

[0032] Cross-modal features are determined based on attention weights and global text features;

[0033] Cross-modal fusion features are determined based on cross-modal features, sequential visual features, global text features, global visual features, and global text representation.

[0034] In one embodiment, cross-modal fusion features are determined based on cross-modal features, sequential visual features, global text features, global visual features, and global text representation, including:

[0035] Based on cross-modal features and sequential visual features, determine the basic fusion features;

[0036] Determine the semantic modulation weights at the word level based on global text features;

[0037] Determine the global semantic modulation weights based on the global text representation;

[0038] Determine the spatial structure modulation weights based on global visual features;

[0039] Cross-modal fusion features are determined based on basic fusion features, word-level semantic modulation weights, global semantic modulation weights, and spatial structure modulation weights.

[0040] In one embodiment, residual fusion features are determined based on cross-modal fusion features and global visual features, including:

[0041] The gating weights are determined based on cross-modal fusion features and global visual features;

[0042] The residual fusion features are determined based on the gating weights, global visual features, and cross-modal fusion features.

[0043] Secondly, this application also provides a device for identifying device status, the device comprising:

[0044] The acquisition module is used to acquire the text lexical units associated with the global visual features and global text features of the target device;

[0045] The selection module is used to select target lexical units from various text lexical units based on global visual features;

[0046] The first determination module is used to determine the global representation of the text based on the target lexical units;

[0047] The second determining module is used to determine the state recognition result of the target device based on global visual features and global text representation.

[0048] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0049] Obtain the text lexical units associated with the global visual features and global text features of the target device;

[0050] Based on global visual features, target lexical units are selected from each text lexical unit.

[0051] Determine the global representation of the text based on the target lexical units;

[0052] The state recognition result of the target device is determined based on global visual features and global text representation.

[0053] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0054] Obtain the text lexical units associated with the global visual features and global text features of the target device;

[0055] Based on global visual features, target lexical units are selected from each text lexical unit.

[0056] Determine the global representation of the text based on the target lexical units;

[0057] The state recognition result of the target device is determined based on global visual features and global text representation.

[0058] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0059] Obtain the text lexical units associated with the global visual features and global text features of the target device;

[0060] Based on global visual features, target lexical units are selected from each text lexical unit.

[0061] Determine the global representation of the text based on the target lexical units;

[0062] The state recognition result of the target device is determined based on global visual features and global text representation.

[0063] The aforementioned device status recognition method and apparatus acquires the text words associated with the global visual features and global text features of the target device. Based on the global visual features, target words are selected from the text words. Based on the target words, a global text representation is determined. Based on the global visual features and the global text representation, the status recognition result of the target device is determined. This application first establishes the association between the global visual features of the target device and each text word, then filters out target words that highly match the current scene based on the global visual features, and constructs an accurate global text representation based on the filtered target words. Finally, the global visual features and the global text representation are fused for status recognition. This method effectively solves the problems of heavy recognition burden, slow response speed, and inaccurate recognition results caused by the large amount of text information and interference from irrelevant redundant information in the background technology. Through a visually guided text word filtering mechanism, invalid text information is eliminated, achieving structured refinement of text semantics, significantly improving the accuracy and efficiency of cross-modal feature alignment, reducing redundant calculations, accelerating model response speed, and strengthening the consistency between image scenes and text semantics, thereby greatly improving the accuracy and robustness of device status recognition in complex power distribution network environments. Attached Figure Description

[0064] Figure 1 This is an application environment diagram of the device status identification method provided in this embodiment;

[0065] Figure 2 This is a flowchart illustrating the first device status identification method provided in this embodiment;

[0066] Figure 3 This is a flowchart illustrating the process of determining target words in this embodiment;

[0067] Figure 4 This is a flowchart illustrating the process of determining the global representation of text provided in this embodiment;

[0068] Figure 5 This is a flowchart illustrating the process of determining the state identification result of the target device in this embodiment.

[0069] Figure 6 This is a flowchart illustrating the second device status identification method provided in this embodiment;

[0070] Figure 7 This is a structural block diagram of a device status identification device provided in this embodiment;

[0071] Figure 8 This is an internal structural diagram of the computer device provided in this embodiment. Detailed Implementation

[0072] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0073] With the accelerated development of smart grids, the distribution network, as a crucial link connecting the transmission network and end users, is receiving increasing attention for its operational reliability, power quality, and safety. Among these, 10kV overhead lines constitute the backbone of the distribution network, widely distributed throughout urban and rural areas, and their operational status directly affects the normal order of social production and daily life.

[0074] Currently, most methods for locating faults and identifying the status of power equipment such as connectors, fuses, insulators, transformers, and overhead switches rely on images or text information fed back from camera equipment, power drone inspections, and autonomous power robot operations. However, current equipment status identification methods based on image or text recognition suffer from significant drawbacks. The large volume of text information not only increases the recognition burden but also results in slow response times and inaccurate identification results.

[0075] To address the aforementioned challenges, the device status identification method provided in this application can be executed by the user-side device, by the server, or through interaction between the user-side device and the server. For example, it can be applied to... Figure 1 In the application environment shown, the server receives the global visual features and global text features of the target device sent by the user-side device. The server acquires each text term associated with the global visual features and global text features of the target device, selects the target term from each text term based on the global visual features, and then determines the global text representation based on the target term. Finally, based on the global visual features and the global text representation, the state recognition result of the target device is determined.

[0076] In this context, "user-side equipment" refers to user terminals, and in this application, it mainly refers to user-side equipment with power equipment inspection functions, such as cameras, inspection robots, and inspection drones. "Server" refers to a server used to receive image and text information sent by the user-side equipment and, based on this information, identify the status of the power equipment. The server can be a standalone server or a server cluster.

[0077] In one embodiment, such as Figure 2 As shown, a device status identification method is provided, which can be applied to... Figure 1 Taking the server-side as an example, the explanation includes the following steps:

[0078] S201, Obtain the text lexical units associated with the global visual features and global text features of the target device.

[0079] Here, "target device" refers to the device that needs to be identified in terms of its status, and can be any device in the power distribution network. "Global visual features" refers to the visual features of the target device extracted from its visual image. "Global textual features" refers to all textual descriptive features related to the target device.

[0080] Optionally, in this embodiment, the image information of the target device sent by the user-side device is received, the image information of the target device is subjected to feature extraction, and after redundancy cleaning, the global visual features of the target device are obtained.

[0081] Another optional implementation involves receiving device description information for the target device sent by the user-side device. Text features are extracted from the device description information to obtain global text features. The global text features are then segmented to obtain the text units associated with each global text feature.

[0082] S202: Select target lexical units from each text lexical unit based on global visual features.

[0083] Among them, target words refer to words in each text that are strongly related to the device status.

[0084] One possible implementation involves constructing a lightweight binary classification network. The network inputs a fused vector of global visual features and single-text word features, and outputs the binary classification probability of the word as "valid" or "invalid." Each text word is sequentially classified and predicted. Words predicted as "valid" are identified as target words, while redundant words predicted as "invalid" are directly filtered out.

[0085] Another optional implementation involves calculating the feature similarity (such as cosine similarity or Euclidean distance) between each text word feature and the global visual features. The higher the similarity value, the higher the semantic matching degree between the word and the current image scene. A fixed similarity threshold is preset, and the similarity results of all text words are compared one by one. Only words with a similarity greater than the threshold are retained as target words, while redundant words with a similarity lower than the threshold are removed.

[0086] S203, Determine the global representation of the text based on the target lexical units.

[0087] One possible implementation is to concatenate the target words to obtain a global representation of the text.

[0088] Another alternative implementation is to treat all the filtered target word features as a feature sequence, and calculate the arithmetic mean of all word features in the sequence dimension by dimension. The mean vector is the global representation of the text.

[0089] S204. Determine the state recognition result of the target device based on global visual features and global text representation.

[0090] One possible implementation involves directly concatenating and fusing global visual features and global text representations along the feature dimension to obtain a joint feature vector that simultaneously contains visual scene information and textual semantic information. This joint feature vector is then input into a classification network consisting of fully connected layers and activation functions. The network learns the correlation mapping relationship between the two types of features and ultimately outputs a state recognition result. This state recognition result may include device state, state category, and confidence level.

[0091] Another alternative implementation involves pre-constructing two parallel inference branches, inputting global visual features and text global representations into independent sub-classification networks, each outputting corresponding device state prediction results; then, using strategies such as voting and weighted summation, the two sets of device state prediction results are comprehensively judged to obtain the final target device state recognition result.

[0092] The aforementioned device status recognition method acquires the text units associated with the global visual features and global text features of the target device. Based on the global visual features, target units are selected from these text units. Based on the target units, a global text representation is determined. Based on the global visual features and the global text representation, the status recognition result of the target device is determined. This application first establishes the association between the global visual features of the target device and each text unit, then filters out target units that highly match the current scene based on the global visual features, and constructs an accurate global text representation based on the filtered target units. Finally, the global visual features and the global text representation are fused for status recognition. This method effectively solves the problems of heavy recognition burden, slow response speed, and inaccurate recognition results caused by the large volume of text information and interference from irrelevant redundant information in the background technology. Through a visually guided text unit filtering mechanism, invalid text information is eliminated, achieving structured refinement of text semantics, significantly improving the accuracy and efficiency of cross-modal feature alignment, reducing redundant calculations, accelerating model response speed, and strengthening the consistency between image scenes and text semantics, thereby greatly improving the accuracy and robustness of device status recognition in complex power distribution network environments.

[0093] In one embodiment, in order to more accurately and efficiently filter target words, such as Figure 3 As shown, one possible implementation method for selecting target lexical units from various text lexical units based on global visual features includes:

[0094] S301 performs global average pooling on the global visual features to obtain a global visual representation.

[0095] In this context, the visual global representation refers to the feature vector obtained after pooling.

[0096] Optionally, in this embodiment, global average pooling is performed on the global visual features based on the spatial dimension to obtain a global visual representation. Specifically, the global visual features can be processed using the following formula to obtain a global visual representation:

[0097]

[0098] Among them, v g V represents the global visual representation; H represents the height of the feature map; and W represents the width of the feature map.

[0099] S302, for each text word, the text word is concatenated with the shown visual global representation to obtain the concatenated feature.

[0100] Optionally, in this embodiment, for each text word, the text word and the shown visual global representation can be concatenated using the following formula to obtain the concatenated features:

[0101]

[0102] in, This represents the i-th text term in the global text feature T; Represents the concatenation feature of the i-th text word; v g This represents the overall visual representation.

[0103] S303. Determine the target score for each text word based on the splicing features and visual global representation corresponding to each text word.

[0104] Optionally, in this embodiment, for each text word, a basic score is obtained based on the concatenation features corresponding to the multilayer perceptron and the text word. A prior score for the text word is determined based on the text prior mapping function and the text word. A target score for the text word is determined based on the basic score and the prior score. In this embodiment, an optional implementation of obtaining the basic score of the text word based on the concatenation features corresponding to the multilayer perceptron and the text word is to input the concatenation features into the multilayer perceptron (MLP) to obtain the basic score of the text word; specifically, the basic score of the text word can be obtained using the following formula:

[0105]

[0106] in, Indicates the basic score; This represents the mapping function in a multilayer perceptron, which consists of fully connected layers and nonlinear activation functions. This represents the concatenation feature of the i-th text word.

[0107] Optionally, in this embodiment, an optional implementation method for determining the prior score of a text word based on the text prior mapping function and the text word is to input the text word into the text prior mapping function to obtain the prior score corresponding to the text word; specifically, it can be determined by the following formula:

[0108]

[0109] Among them, in the formula This represents the prior score of the i-th text word; This represents the prior mapping function.

[0110] Optionally, in this embodiment, the possible implementation of determining the target score of a text word based on the base score and the prior score is as follows: determining the weight coefficient of the prior score; determining the product of the prior score and the weight coefficient; and using the sum of the base score and the product of the text word as the target score of the text word.

[0111] S304, Select target words from each text word based on the target score of each text word.

[0112] Optionally, in this embodiment, the text words are sorted according to the target score to obtain a sorting result. Based on the sorting result, target words are selected from the text words. Specifically, the text words are sorted in descending order according to the target score to obtain a sorting result. Based on the sorting result, a preset number of text words at the top of the sorting are selected as target words.

[0113] In this embodiment, global average pooling is used to condense the overall information of visual features, resulting in a global visual representation that characterizes the entire scene of the image. This global representation is then concatenated with the features of each text word to construct a cross-modal joint feature. The target score for each word is calculated comprehensively based on the global visual representation, and valid target words are selected according to the scores. This approach fully utilizes the overall scene information of the image to guide text semantic selection, accurately quantifies the matching degree between different words and the current inspection image, effectively filters redundant and irrelevant text words, reduces subsequent computation, and solves the problems of low recognition efficiency and insufficient accuracy caused by complex text information in traditional solutions. This improves the rationality of word selection and the overall algorithm efficiency.

[0114] In one embodiment, to more accurately determine the target lexical unit, such as Figure 4 As shown, the optional implementation methods for determining the global representation of text based on the target lexical units include:

[0115] S401, determine the word weight of each target word based on the target score of each target word.

[0116] Optionally, in this embodiment, for each target word, a first ratio of the target score to the temperature coefficient is determined. Based on this first ratio and the first ratios of all target words, the word weight of the target word is determined. Specifically, the word weight of each target word can be determined using the following formula:

[0117]

[0118] in, This represents the lexical weight of the i-th target lexical; This represents the target score for the i-th target word element; τ represents the target score for the j-th target word; τ is the temperature coefficient.

[0119] S402, based on the word weights of each target word, perform a weighted summation of each target word to obtain the global representation of the text.

[0120] Optionally, in this embodiment, the global text identifier can be determined using the following formula:

[0121]

[0122] Among them, t g Indicates the global representation of text; This represents the lexical weight of the i-th target lexical; This represents the i-th target word.

[0123] In this embodiment, the weight of each target word is first determined based on its target score. Then, the features of each target word are aggregated using a weighted summation method to generate a global text representation. This method can adaptively assign weights to each target word, giving higher weights to words highly related to device status recognition while weakening the interference of low-related words. This allows the generated global text representation to accurately focus on key semantic information, preserving the semantic contribution of the target words while refining and denoising the information through weighted aggregation. This improves the semantic relevance and scene adaptability of the global text representation, providing more reliable semantic support for subsequent cross-modal feature fusion and device status recognition.

[0124] In one embodiment, to more accurately determine the state identification result of the target device, such as Figure 5 As shown, an optional implementation method for determining the state recognition result of a target device based on global visual features and global text representation includes:

[0125] S501. Based on global visual features, global text features, and global text representation, determine cross-modal fusion features.

[0126] Optionally, in this embodiment, serialized visual features are determined based on global visual features. Attention weights are determined based on serialized visual features and global text features. Cross-modal features are determined based on attention weights and global text features. Cross-modal fusion features are determined based on cross-modal features, serialized visual features, global text features, global visual features, and global text representation.

[0127] In this embodiment, based on the global visual features, the optional implementation method for serializing visual features is to flatten the global visual features into a sequence form to obtain serialized visual features.

[0128] In this embodiment, an optional implementation for determining attention weights based on sequential visual features and global text features is to determine the query vector based on the product of the sequential visual features and the first learning parameter; that is... Where Q represents the query vector; Represents serialized visual features; This represents the first learning parameter. The key vector is determined based on the product of the global text features and the second learning parameter; that is... Where K represents the key vector; T represents the global text feature; This represents the second learning parameter. The value vector is determined based on the product of the global text features and the third learning parameter; that is... ;in, Represents a value vector; T represents global text features; This represents the third learning parameter. The attention weights are determined using the following formula:

[0129]

[0130] Where A represents the attention weight; Q represents the query vector; d represents the transpose of the key vector; d represents the feature dimension.

[0131] In this embodiment, based on attention weights and global text features, an optional implementation method for determining cross-modal features is to use the product of attention weights and global text features as the cross-modal features.

[0132] In this embodiment, the optional implementation method for determining cross-modal fusion features based on cross-modal features, sequential visual features, global text features, global visual features, and global text representation is as follows: First, determine the basic fusion features based on cross-modal features and sequential visual features. Second, determine the word-level semantic modulation weights based on global text features. Third, determine the global semantic modulation weights based on global text representation. Fourth, determine the spatial structure modulation weights based on global visual features. Finally, determine the cross-modal fusion features based on the basic fusion features, word-level semantic modulation weights, global semantic modulation weights, and spatial structure modulation weights. In this embodiment, the basic fusion features are determined based on cross-modal features, sequential visual features, and the following formula:

[0133]

[0134] Where F represents the basic fusion feature; Indicates cross-modal features; This is the fourth learning parameter, i.e., the parameter to be learned; Represents serialized visual features.

[0135] In this embodiment, an optional implementation method for determining the word-level semantic modulation weight based on global text features is to determine the word-level semantic representation according to the following formula:

[0136]

[0137] in, Let T represent the word-level semantic representation; T represents the global text features. The initial modulation weights are determined based on the word-level semantic representation, the first learnable mapping parameters, and the following formula:

[0138]

[0139] in, Indicates the initial modulation weights; This represents semantic representation at the word level; Indicates learnable mapping parameters; This represents the Sigmoid activation function. Then, the initial modulation weights are extended to the spatial dimension to obtain the word-level semantic modulation weights.

[0140] In this embodiment, an optional implementation method for determining the global semantic modulation weight based on the global text representation is as follows: The initial semantic modulation weight is determined based on the global text representation and the following formula:

[0141]

[0142] in, Indicates the initial semantic modulation weights; Indicates the global representation of text; This represents the second learnable mapping parameter; This represents the Sigmoid activation function. Finally, the initial semantic modulation weights are extended to the spatial dimension to obtain the global semantic modulation weights.

[0143] In this embodiment, an optional implementation method for determining the spatial structure modulation weight based on global visual features is as follows: The spatial structure modulation weight is determined based on global visual features and the following formula:

[0144]

[0145] Where V represents global visual features; Indicates spatial structure modulation weights; This represents the Sigmoid activation function; This indicates that a convolution operation is performed on global visual features.

[0146] In this embodiment, based on the basic fusion features, word-level semantic modulation weights, global semantic modulation weights, and spatial structure modulation weights, an optional implementation method for determining the cross-modal fusion features is to use the product of the basic fusion features, word-level semantic modulation weights, global semantic modulation weights, and spatial structure modulation weights as the cross-modal fusion features; that is, the cross-modal fusion features are determined according to the following formula:

[0147]

[0148] in, F represents the cross-modal fusion feature; F represents the basic fusion feature. Indicates the semantic modulation weights at the word level; Represents the global semantic modulation weights; Indicates spatial structure modulation weights; This indicates element-wise multiplication.

[0149] S502, determine residual fusion features based on cross-modal fusion features and global visual features.

[0150] Optionally, in this embodiment, the gating weights are determined based on cross-modal fusion features and global visual features. The residual fusion features are then determined based on the gating weights, global visual features, and cross-modal fusion features.

[0151] In this embodiment, an optional implementation for determining the gating weight based on cross-modal fusion features and global visual features is as follows: The stitching features are determined based on the cross-modal fusion features, global visual features, and the following formula:

[0152]

[0153] Where U represents the splicing feature; V represents the global visual feature; Indicates cross-modal fusion features; This indicates splicing / merging.

[0154] The gating weights are determined based on the splicing characteristics and the following formula:

[0155]

[0156] Where G represents the gating weight; U represents the concatenated feature; This represents the Sigmoid activation function; This represents the convolution operation; This represents a non-linear activation function.

[0157] In this embodiment, an optional implementation method for determining the residual fusion features based on gating weights, global visual features, and cross-modal fusion features is as follows: The residual fusion features are determined based on gating weights, global visual features, cross-modal fusion features, and the following formula:

[0158]

[0159] Where Y represents the residual fusion feature; V represents the global visual feature; and G represents the gating weight. This indicates cross-modal fusion features.

[0160] S503 determines the state identification result of the target device based on the residual fusion characteristics.

[0161] Optionally, in this embodiment, the residual fusion features are input to the detection head to obtain the final state recognition result of the target device.

[0162] In this embodiment, cross-modal fusion features are generated based on global visual features, global text features, and global text representation, achieving deep alignment and interaction between visual scene information and text semantic information. Then, through residual fusion of the cross-modal fusion features and global visual features, the detailed information of the original visual features is effectively preserved, while enhancing the semantic enhancement effect brought by cross-modal interaction and avoiding the loss of original information caused by direct fusion. Finally, target device status is identified based on the residual fusion features. This approach, through multi-stage cross-modal interaction and residual compensation mechanisms, fully utilizes the refined semantic guidance provided by global text representation while preserving the detailed information of the original visual features, improving the robustness and scene adaptability of feature representation. It effectively solves the problems of information loss and insufficient semantic and visual alignment in traditional cross-modal fusion, significantly improving the accuracy and stability of device status identification in complex power inspection scenarios.

[0163] In one embodiment, such as Figure 6 As shown, another optional implementation of a device status identification method is as follows:

[0164] S601, acquire the text lexical units associated with the global visual features and global text features of the target device.

[0165] S602 performs global average pooling on the global visual features to obtain a global visual representation.

[0166] S603, for each text word, the text word is concatenated with the shown visual global representation to obtain the concatenated feature.

[0167] S604: For each text word, a basic score is obtained based on the concatenation features corresponding to the multilayer perceptron and the text word.

[0168] S605, determine the prior score of the text word based on the text prior mapping function and the text word.

[0169] S606 determines the target score for text lexical units based on the base score and prior score.

[0170] S607, based on the target score, sort the text words to obtain the sorting results.

[0171] S608, Select the target word from each text word based on the sorting results.

[0172] S609, determine the word weight of each target word based on the target score of each target word.

[0173] S610: Based on the word weights of each target word, perform a weighted summation of each target word to obtain the global representation of the text.

[0174] S611, determine the serialized visual features based on the global visual features.

[0175] S612 determines the attention weights based on the serialized visual features and global text features.

[0176] S613 determines cross-modal features based on attention weights and global text features.

[0177] S614, determine the basic fusion features based on cross-modal features and sequential visual features.

[0178] S615 determines the semantic modulation weights at the word level based on global text features.

[0179] S616, Determine the global semantic modulation weights based on the global representation of the text.

[0180] S617 determines the spatial structure modulation weights based on global visual features.

[0181] S618. Based on the basic fusion features, word-level semantic modulation weights, global semantic modulation weights, and spatial structure modulation weights, cross-modal fusion features are determined.

[0182] S619 determines the gating weights based on cross-modal fusion features and global visual features.

[0183] S620 determines residual fusion features based on gating weights, global visual features, and cross-modal fusion features.

[0184] S621, Based on the residual fusion characteristics, determine the state identification result of the target device.

[0185] In this embodiment, the text words associated with the global visual features and global text features of the target device are obtained. Target words are selected from these text words based on the global visual features. A global text representation is determined based on the target words. The state recognition result of the target device is determined based on the global visual features and the global text representation. This application first establishes the association between the global visual features of the target device and each text word, then filters out target words that highly match the current scene based on the global visual features, and constructs an accurate global text representation based on the filtered target words. Finally, the global visual features and the global text representation are fused for state recognition. This method effectively solves the problems of heavy recognition burden, slow response speed, and inaccurate recognition results caused by the large volume of text information and interference from irrelevant redundant information in the background technology. Through a visually guided text word filtering mechanism, invalid text information is eliminated, achieving structured refinement of text semantics, significantly improving the accuracy and efficiency of cross-modal feature alignment, reducing redundant calculations, accelerating model response speed, and strengthening the consistency between image scenes and text semantics, thereby greatly improving the accuracy and robustness of device state recognition in complex power distribution network environments.

[0186] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0187] Based on the same inventive concept, this application also provides a device for implementing the device status identification method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more device status identification device embodiments provided below can be found in the limitations of the device status identification method described above, and will not be repeated here.

[0188] In one embodiment, such as Figure 7 As shown, a device status identification device 1 is provided, comprising: an acquisition module 10, a selection module 20, a first determination module 30, and a second determination module 40, wherein:

[0189] The acquisition module 10 is used to acquire the text lexical units associated with the global visual features and global text features of the target device;

[0190] Selection module 20 is used to select target lexical units from each text lexical unit based on global visual features;

[0191] The first determining module 30 is used to determine the global representation of the text based on the target word;

[0192] The second determining module 40 is used to determine the state recognition result of the target device based on global visual features and global text representation.

[0193] In one embodiment, the selection module 20 is further specifically used for:

[0194] Global average pooling is applied to the global visual features to obtain a global visual representation.

[0195] For each text word, the text word is concatenated with the shown visual global representation to obtain the concatenated feature;

[0196] The target score for each text word is determined based on the splicing features and visual global representation corresponding to each text word.

[0197] Based on the target score of each text word, target words are selected from each text word.

[0198] In one embodiment, the selection module 20 is further specifically used for:

[0199] For each text word, a basic score for the text word is obtained based on the multilayer perceptron and the concatenation features corresponding to the text word.

[0200] Based on the text prior mapping function and text words, determine the prior score of the text words;

[0201] The target score for each word in the text is determined based on the base score and prior score.

[0202] In one embodiment, the selection module 20 is further specifically used for:

[0203] Based on the target score, the text words are sorted to obtain the sorting results;

[0204] Based on the sorting results, target terms are selected from each text terminology.

[0205] In one embodiment, the first determining module 30 is further specifically used for:

[0206] Based on the target score of each target word, determine the word weight of each target word;

[0207] Based on the lexical weights of each target lexical, a weighted sum is calculated for each target lexical to obtain a global representation of the text.

[0208] In one embodiment, the second determining module 40 is further specifically used for:

[0209] Based on global visual features, global text features, and global text representation, determine cross-modal fusion features;

[0210] Based on cross-modal fusion features and global visual features, residual fusion features are determined;

[0211] Based on the residual fusion characteristics, the state recognition result of the target device is determined.

[0212] In one embodiment, the second determining module 40 is further specifically used for:

[0213] Based on global visual features, determine the sequential visual features;

[0214] Attention weights are determined based on sequential visual features and global text features;

[0215] Cross-modal features are determined based on attention weights and global text features;

[0216] Cross-modal fusion features are determined based on cross-modal features, sequential visual features, global text features, global visual features, and global text representation.

[0217] In one embodiment, the second determining module 40 is further specifically used for:

[0218] Based on cross-modal features and sequential visual features, determine the basic fusion features;

[0219] Determine the semantic modulation weights at the word level based on global text features;

[0220] Determine the global semantic modulation weights based on the global text representation;

[0221] Determine the spatial structure modulation weights based on global visual features;

[0222] Cross-modal fusion features are determined based on basic fusion features, word-level semantic modulation weights, global semantic modulation weights, and spatial structure modulation weights.

[0223] In one embodiment, the second determining module 40 is further specifically used for:

[0224] The gating weights are determined based on cross-modal fusion features and global visual features;

[0225] The residual fusion features are determined based on the gating weights, global visual features, and cross-modal fusion features.

[0226] Each module in the aforementioned device status identification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0227] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data related to device status identification. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a device status identification method.

[0228] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0229] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0230] Obtain the text lexical units associated with the global visual features and global text features of the target device;

[0231] Based on global visual features, target lexical units are selected from each text lexical unit.

[0232] Determine the global representation of the text based on the target lexical units;

[0233] The state recognition result of the target device is determined based on global visual features and global text representation.

[0234] In one embodiment, when the processor executes the computer program, it further performs the following steps: selecting target lexical units from each text lexical unit based on global visual features, including:

[0235] Global average pooling is applied to the global visual features to obtain a global visual representation.

[0236] For each text word, the text word is concatenated with the shown visual global representation to obtain the concatenated feature;

[0237] The target score for each text word is determined based on the splicing features and visual global representation corresponding to each text word.

[0238] Based on the target score of each text word, target words are selected from each text word.

[0239] In one embodiment, when the processor executes the computer program, it further performs the following steps: determining the target score for each text word based on the concatenation features and visual global representation corresponding to each text word, including:

[0240] For each text word, a basic score for the text word is obtained based on the multilayer perceptron and the concatenation features corresponding to the text word.

[0241] Based on the text prior mapping function and text words, determine the prior score of the text words;

[0242] The target score for each word in the text is determined based on the base score and prior score.

[0243] In one embodiment, when the processor executes the computer program, it further performs the following steps: selecting target words from the text words based on the target scores of each text word, including:

[0244] Based on the target score, the text words are sorted to obtain the sorting results;

[0245] Based on the sorting results, target terms are selected from each text terminology.

[0246] In one embodiment, when the processor executes the computer program, it further performs the following steps: determining a global representation of the text based on target lexical units, including:

[0247] Based on the target score of each target word, determine the word weight of each target word;

[0248] Based on the lexical weights of each target lexical, a weighted sum is calculated for each target lexical to obtain a global representation of the text.

[0249] In one embodiment, when the processor executes the computer program, it further performs the following steps: determining the state recognition result of the target device based on global visual features and text global representation, including:

[0250] Based on global visual features, global text features, and global text representation, determine cross-modal fusion features;

[0251] Based on cross-modal fusion features and global visual features, residual fusion features are determined;

[0252] Based on the residual fusion characteristics, the state recognition result of the target device is determined.

[0253] In one embodiment, when the processor executes the computer program, it further performs the following steps: determining cross-modal fusion features based on global visual features, global text features, and global text representations, including:

[0254] Based on global visual features, determine the sequential visual features;

[0255] Attention weights are determined based on sequential visual features and global text features;

[0256] Cross-modal features are determined based on attention weights and global text features;

[0257] Cross-modal fusion features are determined based on cross-modal features, sequential visual features, global text features, global visual features, and global text representation.

[0258] In one embodiment, when the processor executes the computer program, it further performs the following steps: determining cross-modal fusion features based on cross-modal features, sequential visual features, global text features, global visual features, and text global representation, including:

[0259] Based on cross-modal features and sequential visual features, determine the basic fusion features;

[0260] Determine the semantic modulation weights at the word level based on global text features;

[0261] Determine the global semantic modulation weights based on the global text representation;

[0262] Determine the spatial structure modulation weights based on global visual features;

[0263] Cross-modal fusion features are determined based on basic fusion features, word-level semantic modulation weights, global semantic modulation weights, and spatial structure modulation weights.

[0264] In one embodiment, when the processor executes the computer program, it further performs the following steps: determining residual fusion features based on cross-modal fusion features and global visual features, including:

[0265] The gating weights are determined based on cross-modal fusion features and global visual features;

[0266] The residual fusion features are determined based on the gating weights, global visual features, and cross-modal fusion features.

[0267] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0268] Obtain the text lexical units associated with the global visual features and global text features of the target device;

[0269] Based on global visual features, target lexical units are selected from each text lexical unit.

[0270] Determine the global representation of the text based on the target lexical units;

[0271] The state recognition result of the target device is determined based on global visual features and global text representation.

[0272] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: selecting target lexical units from various text lexical units based on global visual features, including:

[0273] Global average pooling is applied to the global visual features to obtain a global visual representation.

[0274] For each text word, the text word is concatenated with the shown visual global representation to obtain the concatenated feature;

[0275] The target score for each text word is determined based on the splicing features and visual global representation corresponding to each text word.

[0276] Based on the target score of each text word, target words are selected from each text word.

[0277] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining the target score for each text word based on the concatenation features and visual global representation corresponding to each text word, including:

[0278] For each text word, a basic score for the text word is obtained based on the multilayer perceptron and the concatenation features corresponding to the text word.

[0279] Based on the text prior mapping function and text words, determine the prior score of the text words;

[0280] The target score for each word in the text is determined based on the base score and prior score.

[0281] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: selecting target words from the text words based on the target scores of each text word, including:

[0282] Based on the target score, the text words are sorted to obtain the sorting results;

[0283] Based on the sorting results, target terms are selected from each text terminology.

[0284] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining a global representation of the text based on target lexical units, including:

[0285] Based on the target score of each target word, determine the word weight of each target word;

[0286] Based on the lexical weights of each target lexical, a weighted sum is calculated for each target lexical to obtain a global representation of the text.

[0287] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining the state recognition result of the target device based on global visual features and text global representation, including:

[0288] Based on global visual features, global text features, and global text representation, determine cross-modal fusion features;

[0289] Based on cross-modal fusion features and global visual features, residual fusion features are determined;

[0290] Based on the residual fusion characteristics, the state recognition result of the target device is determined.

[0291] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining cross-modal fusion features based on global visual features, global text features, and global text representations, including:

[0292] Based on global visual features, determine the sequential visual features;

[0293] Attention weights are determined based on sequential visual features and global text features;

[0294] Cross-modal features are determined based on attention weights and global text features;

[0295] Cross-modal fusion features are determined based on cross-modal features, sequential visual features, global text features, global visual features, and global text representation.

[0296] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining cross-modal fusion features based on cross-modal features, sequential visual features, global text features, global visual features, and text global representation, including:

[0297] Based on cross-modal features and sequential visual features, determine the basic fusion features;

[0298] Determine the semantic modulation weights at the word level based on global text features;

[0299] Determine the global semantic modulation weights based on the global text representation;

[0300] Determine the spatial structure modulation weights based on global visual features;

[0301] Cross-modal fusion features are determined based on basic fusion features, word-level semantic modulation weights, global semantic modulation weights, and spatial structure modulation weights.

[0302] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining residual fusion features based on cross-modal fusion features and global visual features, including:

[0303] The gating weights are determined based on cross-modal fusion features and global visual features;

[0304] The residual fusion features are determined based on the gating weights, global visual features, and cross-modal fusion features.

[0305] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0306] Obtain the text lexical units associated with the global visual features and global text features of the target device;

[0307] Based on global visual features, target lexical units are selected from each text lexical unit.

[0308] Determine the global representation of the text based on the target lexical units;

[0309] The state recognition result of the target device is determined based on global visual features and global text representation.

[0310] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: selecting target lexical units from various text lexical units based on global visual features, including:

[0311] Global average pooling is applied to the global visual features to obtain a global visual representation.

[0312] For each text word, the text word is concatenated with the shown visual global representation to obtain the concatenated feature;

[0313] The target score for each text word is determined based on the splicing features and visual global representation corresponding to each text word.

[0314] Based on the target score of each text word, target words are selected from each text word.

[0315] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining the target score for each text word based on the concatenation features and visual global representation corresponding to each text word, including:

[0316] For each text word, a basic score for the text word is obtained based on the multilayer perceptron and the concatenation features corresponding to the text word.

[0317] Based on the text prior mapping function and text words, determine the prior score of the text words;

[0318] The target score for each word in the text is determined based on the base score and prior score.

[0319] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: selecting target words from the text words based on the target scores of each text word, including:

[0320] Based on the target score, the text words are sorted to obtain the sorting results;

[0321] Based on the sorting results, target terms are selected from each text terminology.

[0322] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining a global representation of the text based on target lexical units, including:

[0323] Based on the target score of each target word, determine the word weight of each target word;

[0324] Based on the lexical weights of each target lexical, a weighted sum is calculated for each target lexical to obtain a global representation of the text.

[0325] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining the state recognition result of the target device based on global visual features and text global representation, including:

[0326] Based on global visual features, global text features, and global text representation, determine cross-modal fusion features;

[0327] Based on cross-modal fusion features and global visual features, residual fusion features are determined;

[0328] Based on the residual fusion characteristics, the state recognition result of the target device is determined.

[0329] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining cross-modal fusion features based on global visual features, global text features, and global text representations, including:

[0330] Based on global visual features, determine the sequential visual features;

[0331] Attention weights are determined based on sequential visual features and global text features;

[0332] Cross-modal features are determined based on attention weights and global text features;

[0333] Cross-modal fusion features are determined based on cross-modal features, sequential visual features, global text features, global visual features, and global text representation.

[0334] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining cross-modal fusion features based on cross-modal features, sequential visual features, global text features, global visual features, and text global representation, including:

[0335] Based on cross-modal features and sequential visual features, determine the basic fusion features;

[0336] Determine the semantic modulation weights at the word level based on global text features;

[0337] Determine the global semantic modulation weights based on the global text representation;

[0338] Determine the spatial structure modulation weights based on global visual features;

[0339] Cross-modal fusion features are determined based on basic fusion features, word-level semantic modulation weights, global semantic modulation weights, and spatial structure modulation weights.

[0340] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: determining residual fusion features based on cross-modal fusion features and global visual features, including:

[0341] The gating weights are determined based on cross-modal fusion features and global visual features;

[0342] The residual fusion features are determined based on the gating weights, global visual features, and cross-modal fusion features.

[0343] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0344] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0345] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for identifying device status, characterized in that, The method includes: Obtain the text lexical units associated with the global visual features and global text features of the target device; Based on the global visual features, target lexical units are selected from each text lexical unit; Based on the target lexical units, determine the global representation of the text; The state recognition result of the target device is determined based on the global visual features and the global text representation.

2. The method according to claim 1, characterized in that, The step of selecting target lexical units from each text lexical unit based on the global visual features includes: The global visual features are subjected to global average pooling to obtain a global visual representation; For each text word, the text word is concatenated with the shown visual global representation to obtain the concatenated feature; The target score for each text word is determined based on the splicing features corresponding to each text word and the aforementioned visual global representation. Based on the target score of each text word, target words are selected from each text word.

3. The method according to claim 2, characterized in that, The step of determining the target score for each text word based on the concatenation features corresponding to each text word and the visual global representation includes: For each of the text words, a basic score for the text word is obtained based on the multilayer perceptron and the concatenation features corresponding to the text word. Based on the text prior mapping function and the text lexical, determine the prior score of the text lexical; The target score for the text lexical unit is determined based on the base score and the prior score.

4. The method according to claim 2, characterized in that, The step of selecting target words from each text word based on the target score of each text word includes: Based on the target score, each text word is sorted to obtain the sorting result; Based on the sorting results, target terms are selected from each text terminology.

5. The method according to claim 2, characterized in that, The step of determining the global representation of the text based on the target lexical units includes: Based on the target score of each target word, determine the word weight of each target word; Based on the lexical weights of each target lexical, a weighted sum is calculated for each target lexical to obtain a global representation of the text.

6. The method according to claim 1, characterized in that, The step of determining the state recognition result of the target device based on the global visual features and the global text representation includes: Based on the global visual features, the global text features, and the global text representation, determine the cross-modal fusion features; Based on the cross-modal fusion features and the global visual features, the residual fusion features are determined; Based on the residual fusion features, the state recognition result of the target device is determined.

7. The method according to claim 6, characterized in that, The step of determining cross-modal fusion features based on the global visual features, the global text features, and the global text representation includes: Based on the global visual features, determine the serialized visual features; The attention weights are determined based on the serialized visual features and the global text features. Based on the attention weights and the global text features, determine the cross-modal features; Cross-modal fusion features are determined based on the cross-modal features, the sequential visual features, the global text features, the global visual features, and the global text representation.

8. The method according to claim 7, characterized in that, The step of determining cross-modal fusion features based on the cross-modal features, the sequential visual features, the global text features, the global visual features, and the global text representation includes: Based on the cross-modal features and the sequential visual features, the basic fusion features are determined; Based on the global text features, determine the word-level semantic modulation weights; Determine the global semantic modulation weights based on the global text representation; Based on the global visual features, determine the spatial structure modulation weights; Cross-modal fusion features are determined based on the basic fusion features, the word-level semantic modulation weights, the global semantic modulation weights, and the spatial structure modulation weights.

9. The method according to claim 7, characterized in that, The step of determining residual fusion features based on the cross-modal fusion features and the global visual features includes: The gating weights are determined based on the cross-modal fusion features and the global visual features. The residual fusion features are determined based on the gating weights, the global visual features, and the cross-modal fusion features.

10. A device for identifying equipment status, characterized in that, The device includes: The acquisition module is used to acquire the text lexical units associated with the global visual features and global text features of the target device; The selection module is used to select target lexical units from the text lexical units based on the global visual features; The first determining module is used to determine the global representation of the text based on the target lexical units; The second determining module is used to determine the state recognition result of the target device based on the global visual features and the global text representation.