Abnormal identification method based on target video, electronic device and storage medium
By building a device database and combining a large language model, the problem that multimodal systems are difficult to understand operational processes and safety specifications in specific environments is solved, and the accuracy of abnormal recognition is improved.
Patent Information
- Application Number
- CN202510381121.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing multimodal system is difficult to understand specific operating procedures and safety specifications in specific environments such as chemical production workshops, resulting in low accuracy of abnormal identification results.
By building a device database, using the clustering technology of historical videos and device parameter list groups, rich domain knowledge is generated, and combined with large language models, the feature vectors and related data sets of the target video are input to obtain exception recognition results.
Improve the accuracy of abnormal identification results, and provide more accurate abnormal identification support by comprehensively utilizing a variety of data sources.
Smart Images

Figure CN119888584B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of abnormality recognition, and in particular to an abnormality recognition method based on a target video, an electronic device and a storage medium. Background Art
[0002] In recent years, with the development of natural language processing technology, large language models have been widely used in various fields. They are not limited to text processing. They are also combined with other technologies such as computer vision models to form multimodal systems. Multimodal systems can process various types of data (such as images, videos, texts, vectors, etc.) and combine these data through multimodal fusion technology to perform more complex recognition tasks. For example, anomaly recognition tasks combine the capabilities of large language models and computer vision models, and can comprehensively analyze multimodal data and output anomaly recognition results.
[0003] In the prior art, an image or video corresponding to an area is input into a multimodal system, and an abnormality recognition result output by the multimodal system is obtained. A user determines whether there is an abnormality in the area corresponding to the image or video based on the abnormality recognition result, wherein the abnormality includes but is not limited to safety hazards, operational errors, and the like.
[0004] However, the above method also has the following technical problems:
[0005] Although multimodal systems can provide more information than a single modality, they may still lack sufficient domain knowledge or contextual understanding in some cases, thus affecting the accuracy of judgment. For example, in a chemical production workshop, specific operating procedures and safety regulations may not be easily fully understood and applied by general models. Therefore, today's multimodal systems can only simply determine whether there are abnormalities in the area based on images or videos to generate abnormality recognition results, and the accuracy of the obtained abnormality recognition results is low. Summary of the invention
[0006] In view of the above technical problems, the technical solution adopted by the present invention is:
[0007] According to a first aspect of the present invention, a method for identifying anomalies based on a target video is provided, the method comprising the following steps:
[0008] S1. Whenever a preset duration is reached, the video captured by the target video capture device within the preset duration is taken as SP and MB is obtained, where SP is the target video and MB is the target device parameter list set corresponding to SP.
[0009] S2. Input the feature vector of SP, the feature vector list corresponding to MB, and the total set of related data sets corresponding to MB into the preset large language model to obtain the abnormality recognition result corresponding to SP output by the preset large language model, wherein recall is performed in the device database according to SP and MB to obtain the total set of related data sets corresponding to MB.
[0010] Before step S1, the following steps are included to build a device database:
[0011] S01. Clustering feature vectors of all historical videos corresponding to a designated device to obtain a plurality of first center vectors corresponding to the designated device, wherein the duration of the historical video is the same as the duration of the SP.
[0012] S02. If the feature vector of the historical video is in the first cluster corresponding to the first central vector, the feature vector of the designated device parameter list group corresponding to the historical video is used as the intermediate vector corresponding to the first central vector, and the state label corresponding to the designated device parameter list group is used as the state label corresponding to the intermediate vector.
[0013] S03. Cluster all intermediate vectors corresponding to the first central vector to obtain a second central vector corresponding to the first central vector and a state label corresponding to the second central vector.
[0014] S04. Store the first central vector corresponding to the designated device, the second central vector corresponding to the first central vector, and the state label corresponding to the second central vector in a device database.
[0015] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the aforementioned method.
[0016] According to a third aspect of the present invention, there is provided an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the aforementioned method when executing the computer program.
[0017] The present invention has at least the following beneficial effects:
[0018] The present invention provides an abnormality recognition method based on a target video, an electronic device and a storage medium. The method can obtain a target video and a target device parameter list set corresponding to the target video, input a feature vector of the target video, a feature vector list corresponding to the target device parameter list set, and a total set of related data sets corresponding to the target device parameter list set obtained by recalling the target video and the target device parameter list set in a device database into a preset large prediction model to obtain an abnormality recognition result, wherein a device database is constructed based on historical videos corresponding to a specified device, a specified device parameter list group corresponding to the historical videos, and a status label corresponding to the specified device parameter list group. It can be seen that in the present invention, the device database is based on the specified device corresponding to the historical video. The method is constructed by using historical videos, designated device parameter list groups corresponding to the historical videos, and status labels corresponding to the designated device parameter list groups, which contains rich domain knowledge. The total set of related data sets corresponding to the target device parameter list set is a set recalled from the device database, which can provide necessary data support for the anomaly recognition task. The feature vector of the target video, the feature vector list corresponding to the target device parameter list set, and the total set of related data sets corresponding to the target device parameter list set are input into the preset large prediction model to obtain the anomaly recognition result. It comprehensively utilizes multiple data sources, rather than simply judging whether there is an anomaly in the area based on the image or video to generate the anomaly recognition result, which is beneficial to improving the accuracy of the obtained anomaly recognition results. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 A flowchart of an abnormality recognition method based on a target video is provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar tasks, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0023] An embodiment of the present invention provides a method for identifying anomalies based on a target video, the method comprising the following steps: Figure 1 As shown:
[0024] S1. Whenever the preset duration is reached, the video captured by the target video capture device within the preset duration is taken as SP and MB is obtained, where SP is the target video, and MB is the target device parameter list set corresponding to SP. MB includes a target device parameter list group corresponding to several target devices, and the target device parameter list group includes a target device parameter list corresponding to each parameter of the target device. The target device parameter list includes specific parameter values of the parameters within each second of SP, wherein the preset duration is a duration pre-set by a technician in this field according to actual needs, for example: 1 minute, 2 minutes, which will not be repeated here.
[0025] Specifically, the parameters of a device refer to the specific performance indicators of the device during operation, such as power, speed, accuracy, temperature, etc.
[0026] Specifically, the target video acquisition device is a video acquisition device arranged in a target area. In a specific embodiment, the target area may be a chemical production workshop.
[0027] Furthermore, several video acquisition devices are arranged in the target area.
[0028] Specifically, the target device is a device located in the video acquisition area of the target video acquisition device. In a specific embodiment, the device located in the video acquisition area of the target video acquisition device can be chemical production equipment (for example: reactor, material transmission device, polymerization reactor, distillation tower, extraction tower) and detection equipment (for example: gas concentration detection equipment, temperature detection equipment, pressure detection equipment, noise detection equipment, particulate matter detection equipment).
[0029] Specifically, each target device is equipped with at least one sensor for collecting parameters of the target device in real time. For example, a reaction kettle is equipped with a temperature sensor, a pressure sensor, and a liquid level sensor.
[0030] Specifically, the target video acquisition device is equipped with a data receiving port, and the data receiving port is used to receive parameters transmitted from the target device and sensors configured therein.
[0031] In a specific embodiment, step S1 includes: whenever a specific time point is reached, taking the video captured by the target video capture device between the specific time point and a previous specific time point as SP and acquiring MB.
[0032] S2. Input the feature vector of SP, the feature vector list corresponding to MB, and the total set of related data sets corresponding to MB into the preset large language model to obtain the abnormal recognition result corresponding to SP output by the preset large language model, wherein recall is performed in the device database according to SP and MB to obtain the total set of related data sets corresponding to MB. Those skilled in the art know that any method of obtaining the feature vector of a video in the prior art belongs to the protection scope of the present invention and will not be repeated here.
[0033] Specifically, the feature vector list corresponding to the MB includes the feature vector of each target device parameter list group in the MB.
[0034] Specifically, the feature vector of the target device parameter list group is a vector obtained by concatenating the feature vectors of each target device parameter list in the order of the target parameter lists in the target device parameter list group, wherein the vector values in the feature vectors of the target device parameter lists correspond one-to-one to the parameter values in the target device parameter lists.
[0035] Specifically, the total set of related data sets corresponding to the MB includes the related data sets corresponding to each target device parameter list group in the MB, and the related data sets corresponding to the target device parameter list group include: a related video feature vector list corresponding to the target device parameter list group, a related device feature vector list corresponding to the target device parameter list group, and a status label list corresponding to the related device feature vector list, wherein the related video feature vector list includes several related video feature vectors, the related device feature vector list includes several related device feature vectors, and the status label list includes a status label corresponding to each related device feature vector.
[0036] Specifically, the abnormality identification result corresponding to the SP includes an abnormality judgment label, and the abnormality judgment label includes no abnormality and abnormality. When the abnormality judgment label is abnormality, the abnormality identification result also includes an abnormal device identifier.
[0037] Further, when the abnormality judgment label in the abnormality identification result corresponding to the SP is no abnormality, it means that all target devices corresponding to the SP have no abnormality.
[0038] Further, when the abnormality judgment label in the abnormality identification result corresponding to the SP is abnormality exists, it means that the target device corresponding to the abnormal device identifier has an abnormality, for example: the temperature of the target device is too high and there is a safety hazard.
[0039] Specifically, the status label includes a normal label and an abnormal label.
[0040] Specifically, when the state label corresponding to the relevant device feature vector is a normal label, it means that there is no abnormality in the relevant device feature vector.
[0041] Specifically, when the state label corresponding to the relevant device feature vector is an abnormal label, it indicates that the relevant device feature vector is abnormal, for example, the vector value in the relevant device feature vector is too large.
[0042] In a specific embodiment, the preset large language model is a general large language model, such as GPT and BERT.
[0043] In a specific embodiment, the preset large language model is a special model obtained by technical personnel in this field after optimizing and adjusting the general large language model through technical means such as fine-tuning for anomaly recognition tasks.
[0044] Through the above steps, the feature vector of the target video, the feature vector list corresponding to the target device parameter list set, and the total set of related data sets corresponding to the target device parameter list set are input into the preset large prediction model to obtain the anomaly recognition result. It comprehensively utilizes multiple data sources, rather than simply judging whether there is an anomaly in the area based on the image or video to generate the anomaly recognition result, which is beneficial to improving the accuracy of the obtained anomaly recognition results.
[0045] Specifically, before step S1, the following steps S01-S04 are included to build a device database:
[0046] S01. Clustering the feature vectors of all historical videos corresponding to the designated device to obtain several first center vectors corresponding to the designated device, wherein, based on the vector similarity between the feature vectors of the historical videos corresponding to the designated device, clustering the feature vectors of all historical videos corresponding to the designated device to obtain several first clusters corresponding to the designated device and taking the vector corresponding to the center position of the first cluster as the first center vector corresponding to the designated device, the first cluster corresponding to the designated device includes the feature vectors of several historical videos corresponding to the designated device; it can be understood as: taking the vector similarity between the feature vectors of the historical videos corresponding to the designated device as the distance metric in the preset clustering algorithm, and clustering the feature vectors of all historical videos corresponding to the designated device using the preset clustering algorithm to obtain several first clusters corresponding to the designated device. Those skilled in the art know that the preset clustering algorithm is a clustering algorithm that does not require the number of clusters to be specified in advance and can use vector similarity as a distance metric, such as: hierarchical clustering method. Any method for obtaining vector similarity between vectors in the prior art, such as: all belong to the protection scope of the present invention, such as: cosine similarity, Euclidean distance, which will not be repeated here.
[0047] Specifically, the vector corresponding to the center position of the first cluster is the average value of all eigenvectors in the first cluster.
[0048] Specifically, before step S01, the method also includes the following steps: for each designated device, a number of historical videos corresponding to the designated device and a designated device parameter list group corresponding to each historical video are obtained, the designated device parameter list group includes a designated device parameter list corresponding to each parameter of the designated device, the designated device parameter list includes specific parameter values of the parameters within each second of the historical video, wherein each designated device parameter list group corresponds to a status label.
[0049] Specifically, the designated device is a device located in the target area.
[0050] Specifically, the historical video corresponding to the designated device is a video collected before the current time point that can present the designated device.
[0051] Specifically, the duration of the historical video is the same as the duration of the SP.
[0052] Specifically, when the status label corresponding to the specified device parameter list group is an abnormal label, it means that the specified device parameter list group is inconsistent with the standard parameter list group corresponding to the specified device in the historical video corresponding to the specified device parameter list group, and there may be a fault or improper operation.
[0053] Specifically, when the status label corresponding to the designated device parameter list group is a normal label, it indicates that the designated device parameter list group is consistent with the standard parameter list group corresponding to the designated device in the historical video corresponding to the designated device parameter list group.
[0054] S02. If the feature vector of the historical video is in the first cluster corresponding to the first central vector, the feature vector of the designated device parameter list group corresponding to the historical video is used as the intermediate vector corresponding to the first central vector, and the state label corresponding to the designated device parameter list group is used as the state label corresponding to the intermediate vector.
[0055] Specifically, the feature vector of the specified device parameter list group is a vector obtained by concatenating the feature vectors of each specified device parameter list in the order of the specified device parameter list in the specified device parameter list group, wherein the vector values in the feature vector of the specified device parameter list correspond one-to-one to the parameter values in the specified device parameter list.
[0056] S03, clustering all the intermediate vectors corresponding to the first central vector, obtaining the second central vector corresponding to the first central vector and the state label corresponding to the second central vector, including the following steps S031-S032:
[0057] S031. Based on the vector similarity between the intermediate vectors corresponding to the first central vector, clustering all the intermediate vectors corresponding to the first central vector to obtain a number of second clusters corresponding to the first central vector, wherein the second cluster corresponding to the first central vector includes a number of intermediate vectors corresponding to the first central vector; it can be understood as: taking the vector similarity between the intermediate vectors corresponding to the first central vector as the distance metric in the preset clustering algorithm, and using the preset clustering algorithm to cluster all the intermediate vectors corresponding to the first central vector to obtain a number of second clusters corresponding to the first central vector.
[0058] S032. If the state labels corresponding to the intermediate vectors in the second cluster corresponding to the first center vector are not completely consistent, the second cluster is divided into two sub-clusters according to the state labels, and the second cluster is deleted and the two sub-clusters are used as two new second clusters corresponding to the first center vector; if the state labels corresponding to the intermediate vectors in the second cluster corresponding to the first center vector are completely consistent, the vector corresponding to the center position of the second cluster is used as the second center vector corresponding to the first center vector, and the state label corresponding to any intermediate vector in the second cluster is used as the state label corresponding to the second center vector.
[0059] Specifically, the second cluster is divided into two subclusters according to the state labels, one subcluster includes all the intermediate vectors of the second cluster whose state labels are normal labels, and the other subcluster includes all the intermediate vectors of the second cluster whose state labels are abnormal labels.
[0060] Specifically, the vector corresponding to the center position of the second cluster is the average value of all intermediate vectors in the second cluster.
[0061] Through the above steps, when the state labels corresponding to the intermediate vectors in the second cluster corresponding to the first central vector are not completely consistent, only one state label cannot represent the state labels corresponding to all the intermediate vectors in the second cluster. Therefore, the second cluster is divided into two sub-clusters according to the state label, and the second cluster is deleted and the two sub-clusters are used as two new second clusters corresponding to the first central vector. The state labels corresponding to all the intermediate vectors in the new second clusters are the same. Only one state label can be used to represent the state labels corresponding to all the intermediate vectors in the new second cluster, which is convenient for storage. There is no need to store the state labels corresponding to all the intermediate vectors in the second cluster in the device database, which can effectively reduce redundant information, reduce the amount of data storage, save storage space, make the device database more concise and efficient, and facilitate subsequent query and use.
[0062] S04. Store the first central vector corresponding to the designated device, the second central vector corresponding to the first central vector, and the state label corresponding to the second central vector in a device database.
[0063] Optionally, the device database is initially NULL.
[0064] Through the above steps, the feature vectors of all historical videos corresponding to the designated device are clustered according to the vector similarity between the feature vectors of the historical videos corresponding to the designated device, and the feature vectors of similar historical videos are clustered into the same first cluster, and the first cluster is represented by the first central vector; the feature vectors of the designated device parameter list group corresponding to the historical videos corresponding to the feature vectors in the first cluster are used as the intermediate vectors corresponding to the first central vector corresponding to the first cluster, and the state labels corresponding to the designated device parameter list group are used as the state labels corresponding to the corresponding intermediate vectors; then, according to the vector similarity between the intermediate vectors corresponding to the first central vector, similar intermediate vectors are clustered into the same second cluster to obtain the second central vector corresponding to the first central vector and the state labels corresponding to the second central vector; the first central vector corresponding to the designated device, the second central vector corresponding to the first central vector, and the state labels corresponding to the second central vector are stored in the device database, so that the device database contains rich domain knowledge, can provide data support for the abnormality recognition task, and through multi-level (first cluster, second cluster) clustering analysis, redundant information can be effectively reduced, the amount of data storage can be reduced, and storage space can be saved, so that the device database is more concise and efficient, and is convenient for subsequent query and use.
[0065] Specifically, in step S2, the process of recalling the device database according to the SP and the MB to obtain the total set of related data sets corresponding to the MB, and obtaining the related data sets corresponding to the target device parameter list group according to the SP, the target device parameter list group in the MB, the target device corresponding to the target device parameter list group, and the device database includes the following steps S051-S054:
[0066] S051, obtain a third central vector list A corresponding to a designated device that is the same device as the target device = {A 1 , A 2 , ..., A i , ..., A m}, A i is the i-th third center vector corresponding to the specified device which is the same device as the target device, the value of i ranges from 1 to m, m is the number of third center vectors corresponding to the specified device which is the same device as the target device, and the third center vector corresponding to the specified device is the second center vector corresponding to the first center vector corresponding to the specified device.
[0067] S052. Get B and A i The average vector similarity C between i , where B is the feature vector of the target device parameter list group.
[0068] Specifically, step S052 includes the following steps S0521-S0523:
[0069] S0521. Decompose B to obtain the first sub-vector list D={D 1 , D 2 , ..., D j , ..., D n}, D j is the first subvector corresponding to the j-th parameter of the target device in B, the value of j ranges from 1 to n, and n is the number of parameters of the target device; the first subvector corresponding to the j-th parameter of the target device can be understood as: the feature vector of the target device parameter list corresponding to the j-th parameter of the target device.
[0070] S0522, A i Decompose to obtain the second sub-vector list E i ={E i1 , E i2 ,……,E ij ,……,E in}, E ij A i The second subvector corresponding to the jth parameter of the target device in A; i The second subvector corresponding to the jth parameter of the target device in A can be understood as:i The average value of all the corresponding specified feature vectors in the corresponding second cluster, where the jth parameter of the target device is in A i The corresponding designated feature vector in the corresponding second cluster is a feature vector of the designated device parameter list corresponding to the j-th parameter of the target device in the designated device parameter list group corresponding to the intermediate vector in the second cluster.
[0071] S0523, according to D j and E ij Get C i , where C i Meet the following conditions:
[0072] C i =∑ n j=1 F ij / n,F ij D j With E ij The vector similarity between .
[0073] Specifically, the greater the vector similarity, the more similar the two vectors are.
[0074] Through the above steps, the parameters of the target device are used as the splitting dimensions, the characteristic vector of the target device parameter list group is split into a first sub-vector corresponding to the parameters of the target device, and the third central vector is split into a second sub-vector corresponding to the parameters of the target device. Based on the vector similarity between the first sub-vector and the second sub-vector corresponding to the parameters of the target device, the average vector similarity between the characteristic vector of the target device parameter list group and the third central vector is obtained, which can smooth out the influence of individual abnormal parameter values.
[0075] S053, if C i ≥C 0 , then A i As the relevant device feature vector corresponding to the target device parameter list group, A i The corresponding state label is A i The corresponding state label of the relevant device feature vector is A i The corresponding first center vector is used as the relevant video feature vector corresponding to the target device parameter list group, where C 0 It is a preset vector similarity threshold. Those skilled in the art know that the preset vector similarity threshold is a value less than 1 pre-set by those skilled in the art according to actual needs, for example, 0.8, 0.9, 0.95, which will not be repeated here.
[0076] Through the above steps, the average vector similarity between the feature vector of the target device parameter list group and the third center vector corresponding to the designated device which is the same device as the target device is obtained. If the average similarity between the feature vector of the target device parameter list group and the third center vector is not less than the preset vector similarity threshold, it means that the feature vector of the target device parameter list group is very similar to the third center vector. It can be understood that the target device parameter list group and the designated device parameter list group corresponding to the feature vector in the second cluster corresponding to the third center vector are also similar. Therefore, using the third center vector as the relevant device feature vector corresponding to the target device parameter list group, using the state label corresponding to the third center vector as the state label corresponding to the relevant device feature vector, and using the first center vector corresponding to the third center vector as the relevant video feature vector corresponding to the target device parameter list group can provide necessary data support for the abnormality recognition task, which is conducive to improving the accuracy of the obtained abnormality recognition results.
[0077] S054, if C 1 <C 0 , C 2 <C 0 , ..., C i <C 0 , ..., C m <C 0 , then obtaining the relevant data set corresponding to the target device parameter list group according to SP, including the following steps S0541-S0544:
[0078] S0541. Obtain a target sub-video corresponding to the target device from the SP, wherein the target sub-video is a partial video in the SP that only presents the target device. Those skilled in the art know that any method for obtaining a sub-video from a video in the prior art falls within the protection scope of the present invention and will not be described in detail herein.
[0079] S0542, obtain the first center vector list G corresponding to the designated device which is the same device as the target device={G 1 , G 2 , ..., G e , ..., G f}, where G e is the e-th first center vector corresponding to the designated device that is the same device as the target device, the value of e ranges from 1 to f, and f is the number of first center vectors corresponding to the designated device that is the same device as the target device.
[0080] S0543, obtain the feature vector of the target sub-video and G e The vector similarity H between e .
[0081] S0544, if He ≥C 0 , then G e As the relevant video feature vector corresponding to the target device parameter list group, G e The corresponding second center vector is used as the relevant device feature vector corresponding to the target device parameter list group; if H 1 <C 0 , H 2 <C 0 , ..., H e <C 0 , ..., H f <C 0 , then H 1 , H 2 , ..., H e , ..., H f The first central vector corresponding to the largest vector similarity is used as the relevant video feature vector corresponding to the target device parameter list group, and the second central vector corresponding to the first central vector is used as the relevant device feature vector corresponding to the target device parameter list group, wherein the state label corresponding to the second central vector is used as the state label corresponding to its corresponding relevant device feature vector.
[0082] Through the above steps, if the average similarity between the feature vector of the target device parameter list group and any third central vector is less than the preset vector similarity threshold, it means that the feature vector of the target device parameter list group is not so similar to these third central vectors, and the data related to the target device parameter list group cannot be obtained according to the third central vector. At this time, the target sub-video corresponding to the target device is obtained from the target video, and the vector similarity between the feature vector of the target sub-video and the first central vector corresponding to the designated device which is the same device as the target device is obtained. When the vector similarity is not less than the preset vector similarity threshold, it means that the feature vector of the target sub-video is very similar to the first central vector. It can be understood that the historical video corresponding to the feature vector in the first cluster corresponding to the target sub-video and the first central vector is also similar. Therefore, the first central vector is used as the relevant video feature vector corresponding to the target device parameter list group, the second central vector corresponding to the first central vector is used as the relevant device feature vector corresponding to the target device parameter list group, and the state label corresponding to the second central vector is used as the state label corresponding to the corresponding relevant device feature vector; it can provide necessary data support for the abnormality recognition task. If the vector similarity between the feature vector of the target sub-video and the first central vector corresponding to the designated device which is the same device as the target device is less than the preset vector similarity threshold, it means that the feature vector of the target sub-video is not so similar to these first central vectors. In this case, compared with being unable to provide necessary data support for the abnormality recognition task, taking the first central vector corresponding to the largest vector similarity threshold as the relevant video feature vector corresponding to the target device parameter list group, taking the second central vector corresponding to the first central vector as the relevant device feature vector corresponding to the target device parameter list group, and taking the status label corresponding to the second central vector as the status label corresponding to its corresponding relevant device feature vector is the best choice, which can provide necessary data support for the abnormality recognition task and is conducive to improving the accuracy of the obtained abnormality recognition results.
[0083] The present invention also provides a specific embodiment, in which the relevant data set corresponding to the target device parameter list group includes: a relevant video feature vector list corresponding to the target device parameter list group, a relevant device feature vector list corresponding to the target device parameter list group, and a status label list corresponding to the relevant video feature list, wherein the relevant video feature vector list includes several relevant video feature vectors, the relevant device feature vector list includes several relevant device feature vectors, and the status label list includes a status label corresponding to each relevant video feature vector.
[0084] Before step S1, the following steps S001-S004 are included to build a device database:
[0085] S001. For each designated device, a number of historical videos corresponding to the designated device and a designated device parameter list group corresponding to each historical video are obtained. The designated device parameter list group includes a designated device parameter list corresponding to each parameter of the designated device. The designated device parameter list includes specific parameter values of the parameters within each second of the historical video, wherein each historical video corresponds to a status label and a list of abnormal parameter names.
[0086] Specifically, when the status tag corresponding to the historical video is a normal tag, it means that the sub-video corresponding to the specified device in the historical video is consistent with the standard video corresponding to the specified device parameter list group, and the sub-video corresponding to the specified device is a partial video in the historical video that only presents the specified device.
[0087] Specifically, when the status tag corresponding to the historical video is an abnormal tag, it means that the sub-video corresponding to the specified device in the historical video is inconsistent with the standard video corresponding to the specified device parameter list group, and there may be a malfunction or improper operation.
[0088] Specifically, when the status label corresponding to the historical video is a normal label, the abnormal parameter name list corresponding to the historical video is NULL.
[0089] Specifically, when the status label corresponding to the historical video is an abnormal label, the abnormal parameter name list corresponding to the historical video includes several abnormal parameter names, and the abnormal parameter names are names of parameters that cause the historical video to be abnormal.
[0090] S002. Clustering the feature vectors of all the specified device parameter list groups corresponding to the specified device to obtain several intermediate clusters corresponding to the specified device, and taking the vector corresponding to the center position of the intermediate cluster as the intermediate center vector corresponding to the specified device, wherein, based on the vector similarity between the feature vectors of the specified device parameter list groups corresponding to the specified device, clustering the feature vectors of all the specified device parameter list groups corresponding to the specified device to obtain several intermediate clusters corresponding to the specified device and taking the vector corresponding to the center position of the intermediate cluster as the intermediate center vector corresponding to the specified device, the intermediate cluster corresponding to the specified device includes the feature vectors of several specified device parameter list groups corresponding to the specified device; it can be understood as: taking the vector similarity between the feature vectors of the specified device parameter list groups corresponding to the specified device as the distance metric in the preset clustering algorithm, and using the preset clustering algorithm to cluster the feature vectors of all the specified device parameter list groups corresponding to the specified device to obtain several intermediate clusters corresponding to the specified device.
[0091] Specifically, the vector corresponding to the center position of the middle cluster is the average value of all feature vectors in the middle cluster.
[0092] Specifically, the designated device parameter list group corresponding to the designated device can be understood as the designated device parameter list group corresponding to the historical video corresponding to the designated device.
[0093] S003. If the feature vector of the specified device parameter list group is in the intermediate cluster corresponding to the intermediate center vector, the feature vector of the historical video corresponding to the specified device parameter list group is used as the key vector corresponding to the intermediate center vector, the status label corresponding to the historical video is used as the status label corresponding to the key vector, and the abnormal parameter name list corresponding to the historical video is used as the abnormal parameter name list corresponding to the key vector.
[0094] S004. Store the intermediate center vector corresponding to the designated device, the key vector corresponding to the intermediate center vector, the state label corresponding to the key vector, and the abnormal parameter name list corresponding to the key vector in a device database.
[0095] Through the above steps, based on the vector similarity between the feature vectors of the specified device parameter list group corresponding to the specified device, the feature vectors of all the specified device parameter list groups corresponding to the specified device are clustered to obtain several intermediate clusters corresponding to the specified device and the vector corresponding to the center position of the intermediate cluster is used as the intermediate center vector corresponding to the specified device. If the feature vector of the specified device parameter list group is in the intermediate cluster corresponding to the intermediate center vector, the feature vector of the historical video corresponding to the specified device parameter list group is used as the key vector corresponding to the intermediate center vector, the state label corresponding to the historical video is used as the state label corresponding to the key vector, and the abnormal parameter name list corresponding to the historical video is used as the abnormal parameter name list corresponding to the key vector. The intermediate center vector corresponding to the specified device, the key vector corresponding to the intermediate center vector, the state label corresponding to the key vector, and the abnormal parameter name list corresponding to the key vector are stored in the device database. Redundant information can be reduced and storage space can be saved by clustering. The state label corresponding to the key vector and the abnormal parameter name list corresponding to the key vector are stored in the device database, so that the device database contains rich and detailed domain knowledge, which can provide data support for the abnormal recognition task and is conducive to improving the accuracy of the abnormal recognition results obtained.
[0096] In step S2, the total set of related data sets corresponding to MB is obtained by recalling the device database according to SP and MB, and the related data sets corresponding to the target device parameter list group are obtained according to SP, the target device parameter list group in MB, the target device corresponding to the target device parameter list group, and the device database, including the following steps S0051-S0055:
[0097] S0051, obtain the intermediate center vector list R corresponding to the specified device which is the same device as the target device. 1 , R2 , ..., R g , ..., R h} , R g is the g-th intermediate center vector corresponding to the designated device that is the same device as the target device, where g ranges from 1 to h, and h is the number of intermediate center vectors corresponding to the designated device that is the same device as the target device.
[0098] S0052. Obtain B and R g The first average vector similarity K between g .
[0099] Specifically, step S0052 includes the following steps ac:
[0100] a. Decompose B to obtain the first sub-vector list D = {D 1 , D 2 , ..., D j , ..., D n}.
[0101] b. R g Decompose to obtain the intermediate subvector list Q g ={Q g1 , Q g2 , ..., Q gj , ..., Q gn}, Q gj For R g The intermediate subvector corresponding to the jth parameter of the target device in R g The intermediate subvector corresponding to the jth parameter of the target device in R can be understood as: g The average value of all key feature vectors in the corresponding intermediate cluster, where the jth parameter of the target device is in R g The corresponding key feature vector in the corresponding intermediate cluster is a feature vector of the designated device parameter list corresponding to the j-th parameter of the target device in the designated device parameter list group corresponding to the key vector in the intermediate cluster.
[0102] c. According to D j and Q gj Get K g , where K g Meet the following conditions:
[0103] K g =∑ n j=1 U gj / n,U gj D j With Q gj The vector similarity between .
[0104] Through the above steps, the parameters of the target device are used as the splitting dimensions, the characteristic vector of the target device parameter list group is split into the first sub-vector corresponding to the parameters of the target device, and the intermediate center vector is split into the intermediate sub-vector corresponding to the parameters of the target device. Based on the vector similarity between the first sub-vector and the intermediate sub-vector corresponding to the parameters of the target device, the average vector similarity between the characteristic vector and the intermediate center vector of the target device parameter list group is obtained, which can smooth out the influence of individual abnormal parameter values.
[0105] S0053, if K g ≥C 0 , then R g As the relevant device feature vector corresponding to the target device parameter list group, R g The corresponding key vector is used as the relevant video feature vector corresponding to the target device parameter list group, and the state label corresponding to the key vector is used as the state label corresponding to the corresponding relevant video feature vector.
[0106] S0054, if K 1 <C 0 , K 2 <C 0 , ..., K g <C 0 , ..., K h <C 0 , then based on R g The corresponding key vector corresponding to the abnormal parameter name list obtains B and R g The second average vector similarity L between g .
[0107] Specifically, in step S0054, based on R g The corresponding key vector corresponding to the abnormal parameter name list obtains B and R g The second average vector similarity L between g The method comprises the following steps S0061-0065:
[0108] S0061. Get R g The corresponding key vector corresponds to the abnormal parameter name list set M g ={M g1 , M g2 , ..., M gr , ..., M gs(g)},M gr For R g The list of abnormal parameter names corresponding to the rth key vector. The value of r ranges from 1 to s(g), where s(g) is R g The number of corresponding key vectors.
[0109] S0062, M g1 , M g2 , ..., M gr , ..., M gs(g) The union of R g Corresponding important parameter name list N g ={N g1 , N g2 , ..., N gk , ..., N gt(g)},N gk For R g The corresponding kth important parameter name, k ranges from 1 to t(g), t(g) is R g The number of corresponding important parameter names.
[0110] S0063, according to N gk Get the jth parameter of the target device in R g The first importance weight P in the corresponding intermediate cluster jg , P jg Meet the following conditions:
[0111] When the parameter name of the jth parameter of the target device matches N gk If the same, let P jg =1+N 0 gk , N 0 gk M g1 , M g2 , ..., M gr , ..., M gs(g) Including N gk The number of exception parameter name lists with the same exception parameter name; when the parameter name of the jth parameter of the target device is the same as N g1 , N g2 , ..., N gk , ..., N gt(g) If they are not the same, let P jg =1.
[0112] S0064, P 1g , P 2g , ..., P jg , ..., P ng Perform normalization to obtain P jg The corresponding normalized value, and P jg The corresponding normalized value is taken as the jth parameter of the target device in R g The corresponding second importance weight W in the corresponding middle cluster jg .
[0113] Specifically, the larger the second importance weight is, the more important the corresponding parameter is.
[0114] S0065, according to D j , W jg and Q gj Get L g , L g Meet the following conditions:
[0115] L g =∑ n j=1 (W jg ×U gj ) / ∑ n j=1 W jg .
[0116] Through the above steps, a set of abnormal parameter name lists corresponding to the key vectors corresponding to the intermediate central vector corresponding to the designated device which is the same device as the target device is obtained, and the union of all abnormal parameter name lists in the abnormal parameter name list set is used as the important parameter name list corresponding to the intermediate central vector, and the first importance weights corresponding to the parameters of the target device in the intermediate cluster are obtained according to the parameter names of the parameters of the target device, the important parameter names and the number of abnormal parameter name lists containing the important parameter names, and the first importance weights corresponding to all the parameters of the target device in the intermediate cluster are normalized to obtain the second importance weights corresponding to the parameters of the target device in the intermediate cluster, eliminating the dimensional differences between the weights of different parameters, and based on the second importance weights corresponding to the parameters of the target device in the intermediate cluster, the vector similarity between the first sub-vector corresponding to the parameters of the target device and the intermediate sub-vector is obtained to obtain the second average vector similarity between the vector features of the target device parameter list group and the intermediate central vector, taking into account the importance of the parameters themselves, so that the data set is related according to the second average vector similarity, which is conducive to improving the accuracy of the related data set obtained.
[0117] S0055, if L g ≥C 0 , then R g As the relevant device feature vector corresponding to the target device parameter list group, R g The corresponding key vector is used as the relevant video feature vector corresponding to the target device parameter list group. If L 1 <C 0 , L 2 <C 0 , ..., L g <C 0 , ..., L h <C 0 , then L 1 , L 2, ..., L g , ..., L h The intermediate center vector corresponding to the largest second average vector similarity is used as the relevant device feature vector corresponding to the target device parameter list group, and the key vector corresponding to the intermediate center vector is used as the relevant video feature vector corresponding to the target device parameter list group, wherein the state label corresponding to the key vector is used as the state label corresponding to the corresponding relevant video feature vector.
[0118] Through the above steps, the intermediate center vector corresponding to the designated device that is the same device as the target device is obtained, and the first average vector similarity between the feature vector of the target device parameter list group and the intermediate center vector is obtained. When the first average vector similarity is not less than the preset vector similarity, it means that the feature vector of the target device parameter list group is very similar to the intermediate center vector. It can be understood that the designated device parameter list group corresponding to the feature vector in the intermediate cluster corresponding to the target device parameter list group and the intermediate center vector is also very similar. Therefore, the intermediate center vector is used as the relevant device feature vector corresponding to the target device parameter list group, the key vector corresponding to the intermediate center vector is used as the relevant video feature vector corresponding to the target device parameter list group, and the state label corresponding to the key vector is used as the state label corresponding to the corresponding relevant video feature vector, which can provide necessary data support for the abnormality recognition task; if the first average vector similarity between the feature vector of the target device parameter list group and any intermediate center vector is less than the preset vector similarity threshold, then the second average vector similarity between the feature vector of the target device parameter list group and the intermediate center vector is obtained. When the second average vector similarity is not less than the preset vector similarity, it means that the feature vector of the target device parameter list group is very similar to the intermediate center vector. It can be understood that the target device parameter list group and the designated device parameter list group corresponding to the feature vector in the intermediate cluster corresponding to the intermediate center vector are also very similar. Therefore, the intermediate center vector is used as the relevant device feature vector corresponding to the target device parameter list group, the key vector corresponding to the intermediate center vector is used as the relevant video feature vector corresponding to the target device parameter list group, and the state label corresponding to the key vector is used as the state label corresponding to the corresponding relevant video feature vector; it can provide necessary data support for the abnormal recognition task; otherwise, the intermediate center vector corresponding to the largest second average vector similarity is used as the relevant device feature vector corresponding to the target device parameter list group, the key vector corresponding to the intermediate center vector is used as the relevant video feature vector corresponding to the target device parameter list group, and the state label corresponding to the key vector is used as the state label corresponding to the corresponding relevant video feature vector; it can provide necessary data support for the abnormal recognition task; it is beneficial to improve the accuracy of the abnormal recognition results obtained.
[0119] The present invention also provides a specific embodiment, wherein step S2 includes: inputting SP, MB and the associated data sets corresponding to MB into a preset large language model to obtain the abnormality recognition result corresponding to SP output by the preset large language model, wherein the associated data set corresponding to MB includes an associated data subset corresponding to each target device parameter list group in MB, and the associated data subset corresponding to the target device parameter list group is a data set recalled in the device database according to SP and the target device parameter list group, including: an associated video set corresponding to the target device parameter list group, an associated parameter list group set corresponding to the target device parameter list group and a status label list corresponding to the associated parameter list group set, wherein the associated video set includes several associated videos, the associated parameter list group set includes several associated parameter list groups, and the status label list includes a status label corresponding to each associated parameter list group.
[0120] After step S04, the following step S05 is also included to build a device database:
[0121] S05. Store the historical videos corresponding to each designated device and the designated device parameter list group corresponding to the historical videos in the device database, and establish an association relationship between the first central vector and the historical videos corresponding to the feature vector in the first cluster corresponding to the first central vector, and establish an association relationship between the second central vector and the designated device parameter list group corresponding to the intermediate vector in the second cluster corresponding to the second central vector.
[0122] Step S053 includes: if C i ≥C 0 , then it will be i The specified device parameter list group with an associated relationship is used as the associated parameter list group corresponding to the target device parameter list group, and the state label corresponding to the specified device parameter list group is used as the state label corresponding to the associated parameter list group. i The historical video with the associated relationship corresponding to the first central vector is used as the associated video corresponding to the target device parameter list group.
[0123] Step S0544 includes: if H e ≥C 0 , then it will be combined with G e The historical videos with related relationships are used as the related videos corresponding to the target device parameter list group, and will be combined with G e The designated device parameter list group with an associated relationship with the corresponding second center vector is used as the associated parameter list group corresponding to the target device parameter list group; if H 1 <C 0 , H 2 <C 0 , ..., H e <C 0 , ..., Hf <C 0 , then it will be combined with H 1 , H 2 , ..., H e , ..., H f The historical video associated with the first central vector corresponding to the largest vector similarity is taken as the associated video corresponding to the target device parameter list group, and the specified device parameter list group associated with the second central vector corresponding to the first central vector is taken as the associated parameter list group corresponding to the target device parameter list group, wherein the status label corresponding to the specified device parameter list group is taken as the status label corresponding to the corresponding associated parameter list group.
[0124] Through the above steps, the historical videos corresponding to the designated device and the designated device parameter list groups corresponding to the historical videos are also stored in the device database, and an association relationship is established between the first central vector and the historical videos corresponding to the feature vector in the first cluster corresponding to the first central vector, and an association relationship is established between the second central vector and the designated device parameter list groups corresponding to the intermediate vector in the second cluster corresponding to the second central vector, thereby enriching the content of the device database. According to the target video and the target device parameter list group, the data set recalled in the device database includes: an associated video set corresponding to the target device parameter list group, an associated parameter list group set corresponding to the target device parameter list group, and a status label list corresponding to the associated parameter list group set, wherein the associated video set includes several associated videos, the associated parameter list group set includes several associated parameter list groups, and the status label list includes a status label corresponding to each associated parameter list group, which can provide more comprehensive data support for the abnormality recognition task and is conducive to improving the accuracy of the abnormality recognition results obtained.
[0125] The present invention also provides a specific embodiment, wherein the associated data subset corresponding to the target device parameter list group is a data set recalled from the device database based on the SP and the target device parameter list group, including: an associated video set corresponding to the target device parameter list group, an associated parameter list group set corresponding to the target device parameter list group, and a status label list corresponding to the associated video set, wherein the associated video set includes several associated videos, the associated parameter list group set includes several associated parameter list groups, and the status label list includes a status label corresponding to each associated video.
[0126] After step S004, the following step S0041 is also included to build a device database:
[0127] S0041. Store the historical videos corresponding to each designated device and the designated device parameter list group corresponding to the historical videos in the device database, establish an association relationship between the intermediate center vector and the designated device parameter list group corresponding to the feature vector in the intermediate cluster corresponding to the intermediate center vector, and establish an association relationship between the key vector and the historical video corresponding to the key vector.
[0128] Step S0053 includes: if K g ≥C 0 , then it will be combined with R g The specified device parameter list group with an associated relationship is used as the associated parameter list group corresponding to the target device parameter list group, and will be associated with R g The historical video with the associated relationship of the corresponding key vector is used as the associated video corresponding to the target device parameter list group, and the state label corresponding to the historical video is used as the state label corresponding to the corresponding associated video.
[0129] Step S0055 includes: if L g ≥C 0 , then it will be g The specified device parameter list group with an associated relationship is used as the associated parameter list group corresponding to the target device parameter list group, and will be associated with R g The historical videos with the corresponding key vectors are used as the associated videos corresponding to the target device parameter list group; if L 1 <C 0 , L 2 <C 0 , ..., L g <C 0 , ..., L h <C 0 , then it will be 1 , L 2 , ..., L g , ..., L h The designated device parameter list group that is associated with the intermediate center vector corresponding to the largest second average vector similarity is used as the associated parameter list group corresponding to the target device parameter list group, and the historical video that is associated with the key vector corresponding to the intermediate center vector is used as the associated video corresponding to the target device parameter list group, wherein the status label corresponding to the historical video is used as the status label corresponding to the corresponding associated video.
[0130] Through the above steps, the historical videos corresponding to each designated device and the designated device parameter list group corresponding to the historical videos are stored in the device database, and an association relationship is established between the intermediate center vector and the designated device parameter list group corresponding to the feature vector in the intermediate cluster corresponding to the intermediate center vector, and an association relationship is established between the key vector and the historical video corresponding to the key vector, thereby enriching the content of the device database. According to the target video and the target device parameter list group, the data set recalled in the device database includes: an associated video set corresponding to the target device parameter list group, an associated parameter list group set corresponding to the target device parameter list group, and a status label list corresponding to the associated video set, wherein the associated video set includes several associated videos, the associated parameter list group set includes several associated parameter list groups, and the status label list includes a status label corresponding to each associated video, which can provide more comprehensive data support for the abnormal recognition task and is conducive to improving the accuracy of the abnormal recognition results obtained.
[0131] The present invention also provides a specific embodiment, which includes the following steps S11-S12 after step S1:
[0132] S11. According to the preset text corresponding to the SP and the target video acquisition device, a recall text set corresponding to the SP is searched in the text database, where the recall text set includes several recall texts. Those skilled in the art know that any method of searching for recall information in a database in the prior art belongs to the protection scope of the present invention and will not be described in detail herein.
[0133] Specifically, the recalled text set corresponding to the SP can be understood as a set of all texts in the text database that are related to the SP and the preset text corresponding to the target video acquisition device.
[0134] Specifically, the preset text corresponding to the target video acquisition device includes preset question prompt text, such as: "Please check whether there is any abnormality in the device in the video"; "Please identify whether the operation of each device in the video is normal"; "Please check whether there are any safety hazards in the device in the video."
[0135] In a specific embodiment, the preset text corresponding to the target video acquisition device also includes text presenting information related to the video acquisition area of the target video acquisition device, for example: text presenting the range of the video acquisition area of the target video acquisition device, and text presenting the device name, device attributes, device purpose, and device location within the video acquisition area of the target video acquisition device.
[0136] Specifically, the text database includes relevant texts for each specified device, such as: operation manual, preset risk management plan, safe operation steps, fault records, personnel duty roster, clothing standard manual, etc.
[0137] S12. Input the preset text corresponding to the target video acquisition device, the recalled text set corresponding to SP, the feature vector of SP, the feature vector list corresponding to MB and the total set of related data sets corresponding to MB into the preset large language model to obtain the abnormality recognition result corresponding to SP output by the preset large language model.
[0138] In a specific embodiment, step S12 includes: inputting the preset text corresponding to the target video acquisition device, the recalled text set corresponding to SP, SP, MB and the associated data set corresponding to MB into the preset large language model to obtain the abnormality recognition result corresponding to SP output by the preset large language model.
[0139] Through the above steps, according to the preset text corresponding to the target video and the target video acquisition device, the recall text set corresponding to the target video is searched in the text database, and the preset text corresponding to the target video acquisition device, the recall text set corresponding to the target video, the feature vector of the target video, the feature vector list corresponding to the target device parameter list set and the total set of related data sets corresponding to the target device parameter list set are input into the preset large language model to obtain the abnormality recognition result. In combination with the preset text and the recall text set corresponding to the target video acquisition device, more domain expertise can be introduced to help the large language model better understand and apply the background information of the specific field, provide rich context information for the large language model, and help improve the accuracy of the abnormality recognition results obtained.
[0140] The present invention also provides a specific embodiment, which further includes the following step S10 after step S2:
[0141] S10. Construct a historical abnormality recognition result data combination corresponding to the SP based on the abnormality recognition result corresponding to the SP, and insert the historical abnormality recognition result data combination corresponding to the SP into the historical abnormality recognition result data set corresponding to the target video acquisition device corresponding to the SP, wherein the historical abnormality recognition result data combination includes the abnormality recognition result and the initial device parameter list set corresponding to the abnormality recognition result, and the initial device parameter list set corresponding to the abnormality recognition result is the target device parameter list set used to obtain the abnormality recognition result.
[0142] Specifically, if the abnormality judgment label in the abnormality recognition result in the historical abnormality recognition result data combination is no abnormality, and the time point when the historical abnormality recognition result data combination is inserted into the historical abnormality recognition result data set is closest to the current time point, then the initial device parameter list set in the historical abnormality recognition data combination is used as the key device parameter list set corresponding to its corresponding target video acquisition device at the current time point.
[0143] Through the above steps, after the abnormal recognition result is obtained, a historical abnormal recognition result data combination is constructed based on the abnormal recognition result, and the historical abnormal recognition result data combination is inserted into the historical abnormal recognition result data set corresponding to the target video acquisition device. Furthermore, the key device parameter list set corresponding to the target video acquisition device at the current time point is obtained, which can manage and update the historical abnormal recognition result data set of the device, and dynamically adjust the key device parameter list set based on the latest abnormal recognition result, which is helpful for user management and query.
[0144] After step S1 and before step S2, the following steps S100-S300 are also included:
[0145] S100, obtaining the data similarity XS between MB and the key device parameter list set corresponding to the target video acquisition device at the current time point, wherein those skilled in the art know that any method for obtaining the data similarity between two data sets in the prior art belongs to the protection scope of the present invention and will not be repeated here.
[0146] S200, when XS ≥ XS 0 When the target device parameter list group in SP and MB is used, steps S051-S054 are executed to obtain the relevant data set corresponding to the target device parameter list group according to the target device parameter list group in SP and MB, the target device corresponding to the target device parameter list group and the first device database, XS 0 It is a preset data similarity threshold, wherein those skilled in the art know that the preset data similarity threshold is a value less than 1 pre-set by those skilled in the art according to actual needs, for example: 0.8, 0.58, 0.9, 0.95, which will not be repeated here.
[0147] Specifically, before step S1, the method further includes: executing steps S01-S04 to build a first device database.
[0148] Specifically, the greater the data similarity, the more similar the key device parameter list set corresponding to the MB and the target video acquisition device at the current time point is.
[0149] S300, when XS<XS 0 When the target device parameter list group is obtained, steps S0051-S0055 are executed to obtain the relevant data set corresponding to the target device parameter list group according to the target device parameter list group in SP and MB, the target device corresponding to the target device parameter list group and the second device database.
[0150] Specifically, before step S1, the method further includes: executing steps S001-S004 to build a second device database.
[0151] Through the above steps, the first device database and the second device database are constructed, and the data similarity between the target device parameter list set and the key device parameter list set corresponding to the target video acquisition device at the current time point is obtained. When the data similarity is higher than the preset data similarity threshold, it means that the target device parameter list set and the key device parameter list set corresponding to the target video acquisition device at the current time point are very similar, and there may be no abnormality. S051-S054 is executed to obtain the relevant data set corresponding to the target device parameter list group, and the relevant data set is obtained by combining the target device parameter list set and the target video. Otherwise, it means that the target device parameter list set and the key device parameter list set corresponding to the target video acquisition device at the current time point are not similar, and it is impossible to judge whether there is an abnormality. At this time, S0051-S0055 is executed to obtain the relevant data set corresponding to the target device parameter list group, and the relevant data set is mainly obtained by relying on the target device parameter list set. The method of obtaining the relevant data set can be flexibly selected, which is conducive to improving the accuracy of obtaining the relevant data set and avoiding unnecessary calculation and resource consumption.
[0152] The present invention also provides a specific embodiment, which further includes the following steps S101-S104 after step S1:
[0153] S101. Input SP and MB into a preset large language model to obtain a target object data set output by the preset large language model, wherein the target object data set includes a target object data list corresponding to several target objects, wherein the target object data list includes the personnel type corresponding to the target object, the device type of the intermediate device corresponding to the target object, the device name of the intermediate device corresponding to the target object, the working status of the intermediate device corresponding to the target object, and the initial intention corresponding to the target object, wherein the target object is an object in SP, the intermediate device corresponding to the target object is a target device in SP with the shortest straight-line distance to the target object, and the target object can be understood as a person in SP.
[0154] Specifically, the intent can be understood as the name of the operation performed on the equipment, such as: checking the equipment, opening the cover of the equipment, adding materials to the equipment, performing daily maintenance, checking the status of the equipment, maintaining and recording logs, and leaving after confirming that the equipment is fault-free.
[0155] S102. Obtain a preset intent tree corresponding to the device type of the intermediate device corresponding to the target object. The structure of the preset intent tree has 6 layers. The first-layer nodes represent the device type, the second-layer nodes represent the device name of the specified device corresponding to the device type, the third-layer nodes represent the working status of the specified device, the fourth-layer nodes represent the personnel type related to the working status of the specified device, the fifth-layer nodes represent the original intent corresponding to the personnel type, and the sixth-layer nodes represent the target intent corresponding to the original intent, wherein each fifth-layer node has only one child node; for example: the first-layer nodes are temperature sensors, the second-layer nodes are sensor A, sensor B, and sensor C, the third-layer nodes are normal operation, failure, and waiting for repair, the fourth-layer nodes are equipment maintenance personnel and equipment inspection personnel, the fifth-layer nodes are for performing daily maintenance and checking equipment status, and the sixth-layer nodes are for completing maintenance and recording logs, confirming that the equipment is fault-free, and then leaving.
[0156] S103: Determine the final intent of the target object according to the target object data list and the nodes in the preset intent tree corresponding to the device type of the intermediate device corresponding to the target object.
[0157] Specifically, step S103 includes the following steps S1031-S1035:
[0158] S1031. When the device name of the intermediate device corresponding to the target object is the same as the device name represented by the second-layer node in the preset intent tree, the second-layer node is used as the second-layer key node.
[0159] S1032: When the working state represented by the child node of the second-layer key node is the same as the working state of the intermediate device corresponding to the target object, use the child node as the third-layer key node.
[0160] S1033: When the personnel type represented by the child node of the third-layer key node is the same as the personnel type corresponding to the target object, the child node is used as the fourth-layer key node.
[0161] S1034. When the original intention represented by the child node of the fourth-layer key node is the same as the initial intention corresponding to the target object, the child node is used as the fifth-layer key node.
[0162] S1035. Use the target intent represented by the child nodes of the fifth-layer key nodes as the final intent of the target object corresponding to the target object data list.
[0163] Through the above steps, the data stored in the target object data list is matched one by one with the nodes in the preset intention tree corresponding to the device type of the intermediate device corresponding to the target object to determine the final intention of the target object, which is conducive to improving the accuracy of the determined final intention.
[0164] S104: Input a target object data list corresponding to the target object, a final intention of the target object, and recall information corresponding to the target object into a preset large language model to obtain an abnormality recognition result corresponding to the target object output by the preset large language model.
[0165] In a specific embodiment, step S104 includes: inputting the SP, the target object data list corresponding to the target object, the final intention of the target object and the recall information corresponding to the target object into a preset large language model to obtain the abnormal recognition result corresponding to the target object output by the preset large language model.
[0166] Specifically, the abnormality recognition results corresponding to the target object include normal and abnormal.
[0167] Further, when the abnormality identification result corresponding to the target object is normal, it means that there is no abnormality in the operation of the target object in the SP.
[0168] Furthermore, when the abnormality identification result corresponding to the target object is abnormal, it indicates that the operation of the target object in the SP is abnormal, which may pose a security risk.
[0169] Specifically, the recall information corresponding to the target object includes video recall information, which is a standard operation video searched from a device database based on the target object data list corresponding to the target object and the ultimate intention of the target object, wherein the device database includes several standard operation videos corresponding to each designated device, and each standard operation video corresponds to a real intention. When the intermediate device corresponding to the target object is the same device as the designated device, and the ultimate intention of the target object is the same as the real intention corresponding to the standard operation video corresponding to the designated device, the standard operation video is used as video recall information.
[0170] In a specific embodiment, the recall information corresponding to the target object includes text recall information, which is a standard operation manual searched from a text database based on the target object data list corresponding to the target object and the ultimate intention of the target object, wherein the text database includes several standard operation manuals corresponding to each designated device, and each standard operation manual corresponds to a real intention. When the intermediate device corresponding to the target object is the same device as the designated device, and the ultimate intention of the target object is the same as the real intention corresponding to the standard operation manual corresponding to the designated device, the standard operation manual is used as text recall information.
[0171] Through the above steps, the target video and the target device parameter list set are input into the preset large language model to obtain the target object data set, and the final intention of the target object is determined according to the preset intention tree corresponding to the device type of the intermediate device corresponding to the target object. The target object data list corresponding to the target object, the final intention of the target object and the recall information corresponding to the target object are input into the preset large language model to obtain the abnormal recognition result corresponding to the target object output by the preset large language model. The recall information corresponding to the target object can provide necessary context information or data support for the abnormal recognition task, and the target object data list corresponding to the target object, the final intention of the target object and the recall information corresponding to the target object are input into the preset large language model to obtain the abnormal recognition result. A variety of data sources are comprehensively utilized, rather than simply judging whether the target object is abnormal based on the image or video to generate the abnormal recognition result, which is conducive to improving the accuracy of the abnormal recognition result obtained.
[0172] The present invention also provides a specific embodiment, which includes the following steps S010-S020 after step S101 and before step S104 to determine the final intention of the target object:
[0173] S010. Obtain a preset knowledge graph, which includes several triples. The first entity in the triple is a preset device type, the second entity is a target intent, and the relationship is a combination of relevant data corresponding to the preset device type and the target intent, including: the device name of the specified device corresponding to the preset device type, the working status of the specified device, the personnel type related to the working status, and the original intent corresponding to the personnel type. For example, the first entity is a temperature sensor, and the second entity is completing maintenance and recording logs. The relationship is a combination of relevant data corresponding to the temperature sensor and completing maintenance and recording logs, including sensor A, fault, equipment maintenance personnel, and performing daily maintenance.
[0174] Optionally, the preset device type is a device type obtained by deduplicating device types of all specified devices.
[0175] Specifically, in a relevant data combination, the number of equipment name, working status, personnel type and original intention are all 1.
[0176] S020. If in the target object data list, the device type of the intermediate device corresponding to the target object is the same as the first entity in the triple, and the intermediate data combination corresponding to the target object data list is completely consistent with the relationship in the triple, then the second entity in the triple is taken as the final intention of the target object, wherein the intermediate data combination corresponding to the target object data list includes: the personnel type corresponding to the target object, the device name of the intermediate device corresponding to the target object, the working status of the intermediate device corresponding to the target object and the initial intention corresponding to the target object.
[0177] Specifically, the intermediate data combination corresponding to the target object data list is completely consistent with the relationship in the triplet, which can be understood as: the personnel type corresponding to the target object in the intermediate data combination is the same as the personnel type related to the working status in the relationship, the device name of the intermediate device corresponding to the target object in the intermediate data combination is the same as the device name of the designated device corresponding to the preset device type in the relationship, the working status of the intermediate device corresponding to the target object in the intermediate data combination is the same as the working status of the designated device in the relationship, and the initial intention corresponding to the target object in the intermediate data combination is the same as the original intention corresponding to the personnel type in the relationship.
[0178] Through the above steps, the structured knowledge graph helps to improve the efficiency and accuracy of data analysis. According to the target object data list and the entities and relationships in the preset knowledge graph, the ultimate intention of the target object can be determined, which can quickly achieve accurate matching and identification of the ultimate intention of the target object, and improve the efficiency of determining the ultimate intention of the target object.
[0179] Specifically, the videos in the above embodiments (target video, historical video, target sub-video, associated video, standard operation video) can be replaced by their corresponding image groups, and the image group includes one frame of image every second of the corresponding video.
[0180] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, which can be set in an electronic device to store a computer program related to a method in a method embodiment, and the computer program is loaded and executed by the processor to implement the method provided in the above embodiment.
[0181] An embodiment of the present invention further provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method provided in the above embodiment when executing the computer program.
[0182] An embodiment of the present invention further provides a computer program product, which includes program code. When the program product is run on an electronic device, the program code is used to enable the electronic device to execute the steps of the method according to various exemplary embodiments of the present invention described above in this specification.
[0183] The present invention provides an abnormality recognition method based on a target video, an electronic device and a storage medium. The method can obtain a target video and a target device parameter list set corresponding to the target video, input a feature vector of the target video, a feature vector list corresponding to the target device parameter list set, and a total set of related data sets corresponding to the target device parameter list set obtained by recalling the target video and the target device parameter list set in a device database into a preset large prediction model to obtain an abnormality recognition result, wherein a device database is constructed based on historical videos corresponding to a specified device, a specified device parameter list group corresponding to the historical videos, and a status label corresponding to the specified device parameter list group. It can be seen that in the present invention, the device database is based on the specified device corresponding to the historical video. The method is constructed by using historical videos, designated device parameter list groups corresponding to the historical videos, and status labels corresponding to the designated device parameter list groups, which contains rich domain knowledge. The total set of related data sets corresponding to the target device parameter list set is a set recalled from the device database, which can provide necessary data support for the anomaly recognition task. The feature vector of the target video, the feature vector list corresponding to the target device parameter list set, and the total set of related data sets corresponding to the target device parameter list set are input into the preset large prediction model to obtain the anomaly recognition result. It comprehensively utilizes multiple data sources, rather than simply judging whether there is an anomaly in the area based on the image or video to generate the anomaly recognition result, which is beneficial to improving the accuracy of the obtained anomaly recognition results.
[0184] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are only for illustration, not for limiting the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention.
Claims
1. A method for identifying anomalies based on a target video, characterized in that: The method comprises the following steps: S1. Whenever a preset duration is reached, the video captured by the target video capture device within the preset duration is used as SP and MB is obtained, where SP is the target video and MB is the target device parameter list set corresponding to SP; S2, inputting the feature vector of SP, the feature vector list corresponding to MB and the total set of related data sets corresponding to MB into a preset large language model to obtain an abnormality recognition result corresponding to SP output by the preset large language model, wherein recall is performed in the device database according to SP and MB to obtain the total set of related data sets corresponding to MB; Before S1, the following steps are included to build the device database: S01, clustering the feature vectors of all historical videos corresponding to the designated device to obtain a plurality of first center vectors corresponding to the designated device, wherein the duration of the historical video is the same as the duration of the SP; S02. If the feature vector of the historical video is in the first cluster corresponding to the first central vector, the feature vector of the designated device parameter list group corresponding to the historical video is used as the intermediate vector corresponding to the first central vector, and the state label corresponding to the designated device parameter list group is used as the state label corresponding to the intermediate vector; S03, clustering all intermediate vectors corresponding to the first central vector, obtaining a second central vector corresponding to the first central vector and a state label corresponding to the second central vector; S04. Store the first central vector corresponding to the designated device, the second central vector corresponding to the first central vector, and the state label corresponding to the second central vector in a device database.
2. The method for identifying anomalies based on target video according to claim 1, characterized in that: The MB includes a target device parameter list group corresponding to several target devices. The target device parameter list group includes a target device parameter list corresponding to each parameter of the target device. The target device parameter list includes specific parameter values of the parameters in each second of the SP.
3. The method for identifying anomalies based on target video according to claim 2, characterized in that: The feature vector list corresponding to MB includes the feature vector of each target device parameter list group in MB. The feature vector of the target device parameter list group is a vector obtained by concatenating the feature vectors of each target device parameter list according to the order of the target parameter lists in the target device parameter list group, wherein the vector values in the feature vectors of the target device parameter list correspond one-to-one to the parameter values in the target device parameter list.
4. The method for identifying anomalies based on target video according to claim 3, characterized in that: The total set of related data sets corresponding to MB includes related data sets corresponding to each target device parameter list group in MB, including: a related video feature vector list corresponding to the target device parameter list group, a related device feature vector list corresponding to the target device parameter list group, and a status label list corresponding to the related device feature vector list, wherein the related video feature vector list includes several related video feature vectors, the related device feature vector list includes several related device feature vectors, and the status label list includes a status label corresponding to each related device feature vector.
5. The method for identifying anomalies based on target video according to claim 4, characterized in that: In step S2, the process of recalling the device database according to the SP and the MB to obtain the total set of related data sets corresponding to the MB, and obtaining the related data sets corresponding to the target device parameter list group according to the SP, the target device parameter list group in the MB, the target device corresponding to the target device parameter list group, and the device database includes the following steps: S051, obtain a third central vector list A={A1, A2, . . . , A ... i , ..., A m }, A i is the i-th third center vector corresponding to the designated device which is the same device as the target device, i ranges from 1 to m, m is the number of third center vectors corresponding to the designated device which is the same device as the target device, and the third center vector corresponding to the designated device is the second center vector corresponding to the first center vector corresponding to the designated device; S052. Get B and A i The average vector similarity C between i , where B is the characteristic vector of the target device parameter list group; S053, if C i ≥C 0 , then A i As the relevant device feature vector corresponding to the target device parameter list group, A i The corresponding state label is A i The corresponding state label of the relevant device feature vector is A i The corresponding first center vector is used as the relevant video feature vector corresponding to the target device parameter list group, where C 0 is the preset vector similarity threshold.
6. The method for identifying anomalies based on target video according to claim 5, characterized in that: Step S052 includes the following steps S0521-S0523: S0521. Decompose B to obtain the first sub-vector list D={D1, D2, ..., D j , ..., D n }, D j is the first subvector corresponding to the jth parameter of the target device in B, where j ranges from 1 to n, and n is the number of parameters of the target device; S0522, A i Decompose to obtain the second sub-vector list E i ={E i1 , E i2 ,……,E ij ,……,E in }, E ij A i The second sub-vector corresponding to the j-th parameter of the target device in ; S0523, according to D j and E ij Get C i , where C i Meet the following conditions: C i =∑ n j=1 F ij / n,F ij D j With E ij The vector similarity between .
7. The method for identifying anomalies based on target video according to claim 5, characterized in that: The step of obtaining the relevant data set corresponding to the target device parameter list group according to the target device parameter list group in the SP and MB, the target device corresponding to the target device parameter list group, and the device database further includes the following steps: S054, if C1<C 0 , C2<C 0 , ..., C i <C 0 , ..., C m <C 0 , then obtain the relevant data set corresponding to the target device parameter list group according to SP.
8. The method for identifying anomalies based on target video according to claim 7, characterized in that: Step S054 includes the following steps S0541-S0544: S0541. Obtain a target sub-video corresponding to the target device from the SP, wherein the target sub-video is a portion of the video in the SP that only presents the target device; S0542, obtain a first central vector list G={G1, G2, . . . , G e , ..., G f }, where G e is the e-th first center vector corresponding to the designated device that is the same device as the target device, where the value of e ranges from 1 to f, and f is the number of first center vectors corresponding to the designated device that is the same device as the target device; S0543, obtain the feature vector of the target sub-video and G e The vector similarity H between e ; S0544, if H e ≥C 0 , then G e As the relevant video feature vector corresponding to the target device parameter list group, G e The corresponding second center vector is used as the relevant device feature vector corresponding to the target device parameter list group; if H1<C 0 , H2<C 0 , ..., H e <C 0 , ..., H f <C 0 , then H1, H2, ..., H e , ..., H f The first central vector corresponding to the largest vector similarity is used as the relevant video feature vector corresponding to the target device parameter list group, and the second central vector corresponding to the first central vector is used as the relevant device feature vector corresponding to the target device parameter list group, wherein the state label corresponding to the second central vector is used as the state label corresponding to its corresponding relevant device feature vector.
9. A non-transitory computer-readable storage medium, characterized in that: The storage medium stores a computer program, which is loaded and executed by a processor to implement the target video-based anomaly recognition method as described in any one of claims 1 to 8.
10. An electronic device comprising: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for identifying anomalies based on a target video as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Content subject discovery method based on multi-modal abnormal content understanding
CN118536049A
Automatic dubbing method and system based on generative AI
CN119314488A