A power service video retrieval method, an electronic device, and a storage medium

CN122885243APending Publication Date: 2026-10-09NANJING SUYI IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611075914.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0004]本发明提供了一种电力服务视频检索方法、电子设备及存储介质,以解决现有电力服务视频检索中因时序语义割裂导致检索精度低、效率差的问题

Benefits of technology

[0010]本发明实施例的技术方案,通过对获取的文本查询进行语义特征提取与分段式交叉变换,得到查询特征;从预设语义特征数据库读取每个待检索视频对应的视频分段特征集合;对于各待检索视频的每个视频分段特征,采用时序对齐匹配算法确定视频分段特征与查询特征之间的对齐距离,并根据对齐距离确定视频分段特征与查询特征的分段相似度;对同一待检索视频的所有分段相似度进行加权聚合,得到待检索视频与文本查询的整体相似度;根据整体相似度对各待检索视频进行排序,并输出检索结果。本方案通过将文本查询进行与视频侧一致的分段式交叉变换,并利用预设语义特征数据库中已存储的视频分段特征集合,结合时序对齐匹配与加权聚合策略,能够实现对电力服务视频的语义级快速检索。该方案避免了在线重复提取视频特征的繁重计算,同时通过跨段语义关联增强了特征表达,通过柔性时序对齐克服了查询与视频长度不一致的问题,从而显著提升了检索精度与响应速度,有效区分不同语义类别的视频内容,克服了传统方法中语义割裂、精度不足、效率低下的缺陷。此外,本方案能高效支撑电网运维故障溯源、操作培训等场景,为智能电网安全运行提供数据支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122885243A_ABST
    Figure CN122885243A_ABST
Patent Text Reader

Abstract

The application discloses a power service video retrieval method, an electronic device and a storage medium. The method comprises the following steps: performing semantic feature extraction and segmented cross transformation on an obtained text query to obtain query features; reading a video segment feature set of each video to be retrieved from a preset semantic feature database; for each video segment feature of each video to be retrieved, a time sequence alignment matching algorithm is used to determine an alignment distance between the video segment feature and the query feature, and the segment similarity between the video segment feature and the query feature is determined according to the alignment distance; all segment similarities of the same video to be retrieved are weighted and aggregated to obtain the overall similarity between the video to be retrieved and the text query; the videos to be retrieved are sorted according to the overall similarity, and a retrieval result is output. The scheme can improve the accuracy and efficiency of power service video retrieval, and provides data support for safe operation of a smart grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, electronic device and storage medium for retrieving power service videos. Background Technology

[0002] With the deepening of smart grid construction, the power service sector has accumulated massive amounts of video data, covering core scenarios such as equipment inspection, fault repair, operation training, and emergency response. These videos not only record key information such as the operating status of power equipment and personnel operating procedures, but are also important data assets for power grid safe operation and maintenance, accident tracing, and skills transfer.

[0003] Existing methods for retrieving power service videos often rely on manually labeled text tags or shallow visual features (such as color and texture), making it difficult to capture deep temporal semantic relationships in videos, such as the evolution of equipment failures and the progression of operational processes. In addition, existing methods either cannot handle global temporal dependencies or have excessively high computational complexity, making it difficult to balance retrieval accuracy and real-time performance, resulting in long retrieval times and insufficient recall in real-world scenarios. Summary of the Invention

[0004] This invention provides a method, electronic device, and storage medium for retrieving power service videos, in order to solve the problems of low retrieval accuracy and poor efficiency caused by temporal semantic fragmentation in existing power service video retrieval.

[0005] According to one aspect of the present invention, a method for retrieving power service videos is provided, the method comprising: Semantic feature extraction and segmented cross-transformation are performed on the acquired text query to obtain query features; Read the video segment feature set corresponding to each video to be retrieved from the preset semantic feature database; For each video segment feature of each video to be retrieved, a temporal alignment matching algorithm is used to determine the alignment distance between the video segment feature and the query feature, and the segment similarity between the video segment feature and the query feature is determined based on the alignment distance. The weighted aggregation of all segment similarities of the same video to be retrieved yields the overall similarity between the video to be retrieved and the text query. The videos to be searched are sorted according to their overall similarity, and the search results are output.

[0006] According to another aspect of the present invention, a power service video retrieval device is provided, the device comprising: The query feature determination module is used to extract semantic features and perform segmented cross-transformation on the acquired text query to obtain query features; The video segmentation feature acquisition module is used to read the set of video segmentation features corresponding to each video to be retrieved from the preset semantic feature database. The matching module is used to determine the alignment distance between the video segment features and the query features for each video segment feature of each video to be retrieved using a temporal alignment matching algorithm, and to determine the segment similarity between the video segment features and the query features based on the alignment distance; The overall similarity determination module is used to perform weighted aggregation of the similarity of all segments of the same video to be retrieved, so as to obtain the overall similarity between the video to be retrieved and the text query. The search results output module is used to sort each video to be searched according to the overall similarity and output the search results.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the power service video retrieval method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the power service video retrieval method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the power service video retrieval method according to any embodiment of the present invention.

[0010] The technical solution of this invention involves extracting semantic features and performing segmented cross-transformation on the acquired text query to obtain query features; reading the video segment feature set corresponding to each video to be retrieved from a preset semantic feature database; for each video segment feature of each video to be retrieved, using a temporal alignment matching algorithm to determine the alignment distance between the video segment feature and the query feature, and determining the segment similarity between the video segment feature and the query feature based on the alignment distance; performing weighted aggregation on all segment similarities of the same video to be retrieved to obtain the overall similarity between the video to be retrieved and the text query; sorting each video to be retrieved based on the overall similarity, and outputting the retrieval results. This solution, by performing segmented cross-transformation on the text query consistent with the video side, and utilizing the video segment feature set already stored in the preset semantic feature database, combined with temporal alignment matching and weighted aggregation strategies, enables rapid semantic-level retrieval of power service videos. This solution avoids the cumbersome computation of repeatedly extracting video features online. It enhances feature representation through cross-segment semantic association and overcomes the problem of inconsistent query and video lengths through flexible temporal alignment, thus significantly improving retrieval accuracy and response speed. It effectively distinguishes video content of different semantic categories, overcoming the shortcomings of traditional methods such as semantic fragmentation, insufficient accuracy, and low efficiency. Furthermore, this solution can efficiently support scenarios such as power grid operation and maintenance fault tracing and operation training, providing data support for the safe operation of smart grids.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a power service video retrieval method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the query feature acquisition process provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the adaptive segmentation process provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the segmented crossover process provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the adaptive segmentation effect provided in an embodiment of the present invention; Figure 6 This is a heatmap of the cross-segment covariance matrix provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the DTW matching distance matrix and optimal path provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the segmentation similarity results provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the retrieval and sorting results provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of a power service video retrieval device provided in an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Figure 1 This is a flowchart illustrating a method for retrieving power service videos according to an embodiment of the present invention. This embodiment is applicable to situations requiring fast and accurate retrieval of power service videos based on text semantics. The method can be executed by a power service video retrieval device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown in the figure, the power service video retrieval method provided in this embodiment specifically includes the following steps: S110. Extract semantic features and perform segmented cross-transformation on the obtained text query to obtain query features.

[0017] Here, text query refers to a natural language statement entered by a user to retrieve power service videos, expressing the semantic needs of the user's desired video content. Semantic feature extraction refers to the process of converting discrete text word sequences into continuous numerical feature vectors, which can capture the deep semantic information and contextual relationships of the text. Segmented cross-transformation refers to a feature transformation method that simultaneously integrates intra-segment temporal information and cross-segment semantic association information. It generates enhanced features that combine intra-segment temporal information and cross-segment association through temporal encoding, cross-segment association extraction, and feature fusion. Query features refer to feature representations obtained after semantic feature extraction and segmented cross-transformation, which have the same dimension as the video segment features and a unified semantic space. These features are used for subsequent temporal alignment matching and similarity calculation with the features of each video segment.

[0018] In this embodiment of the invention, a natural language query statement, such as "transformer fault detection," input by a user can be received. The query statement is then preprocessed, including word segmentation and stop word removal. A pre-trained language model, such as Bidirectional Encoder Representations from Transformers (BERT), is used to extract semantic features from the preprocessed word sequence, yielding the corresponding text semantic features. Since the text semantic features are in single-vector form, while the video segmentation features are in matrix form with temporal dimensions, their dimensions do not match, making direct similarity calculation impossible. Therefore, temporal extension processing is required for the text semantic features. For example, they can be repeated several times along the temporal dimension according to a specified temporal length to obtain the corresponding extended semantic matrix, thus enabling the text representation to be temporally aligned with the video segmentation features. Next, the extended semantic matrix can undergo the same segmented cross-transformation operation as the video segmentation, generating query features with consistent dimensions and a unified semantic space, laying the foundation for subsequent similarity matching with the video segmentation features.

[0019] S120. Read the video segment feature set corresponding to each video to be retrieved from the preset semantic feature database.

[0020] The video to be retrieved can refer to all power service videos stored in a pre-defined power service video library that are available for semantic retrieval by users. Examples include equipment inspection videos, fault repair videos, operation training videos, and emergency response videos. The video segment feature set refers to the set of features corresponding to each video segment after a video to be retrieved has been divided into several semantically unified video segments. Each video segment feature integrates intra-segment temporal information and cross-segment semantic association information for the corresponding video segment. Video segment features can be obtained by performing segmented cross-transformation on the video segments. Intra-segment temporal information refers to the temporal order relationship and continuous content evolution information between different frames within a video segment, reflecting the dynamic change process of the video content. Cross-segment semantic association information refers to the mutual association and logical dependency relationship in semantic content between different video segments, reflecting the global semantic structure of the entire video.

[0021] In this embodiment of the invention, after a user inputs a text query, the retrieval system can access a pre-configured semantic feature database based on query conditions, such as the video retrieval range set by the user. The database retrieves the video segment feature set corresponding to each video to be retrieved, which serves as the basis for subsequent matching calculations. This video segment set can be pre-obtained offline by performing frame sampling, adaptive segmentation, and segmented cross-transformation on the video to be retrieved. The set contains multiple video segment features that correspond one-to-one with each video segment, and each video segment feature integrates intra-segment temporal information and cross-segment semantic association information of the corresponding video segment.

[0022] This embodiment obtains the video segment feature set of each video to be retrieved in the offline stage, so that feature extraction and segmentation processing do not need to be repeated during online retrieval, thereby effectively reducing retrieval latency and improving response speed.

[0023] S130. For each video segment feature of each video to be retrieved, a temporal alignment matching algorithm is used to determine the alignment distance between the video segment feature and the query feature, and the segment similarity between the video segment feature and the query feature is determined based on the alignment distance.

[0024] Temporal alignment matching algorithms refer to a class of algorithms capable of finding the optimal semantic correspondence between two temporal sequences of different lengths. Examples include Dynamic Time Warping (DTW), Hidden Markov Model (HMM) matching algorithms, and attention-based alignment algorithms. Alignment distance refers to the cumulative difference between two feature sequences calculated by a temporal alignment matching algorithm under the optimal temporal alignment path; a smaller value indicates greater similarity between the two sequences. Segment similarity refers to the semantic similarity between video segment features and query features, with a value ranging from [0,1]. A value closer to 1 indicates higher semantic relevance.

[0025] In this embodiment of the invention, for each video segment feature in each video to be retrieved, the best matching distance, i.e., the alignment distance, between the video segment feature and the query feature can be calculated using temporal alignment algorithms such as DTW. Then, the distance is converted into the semantic similarity between the video segment feature and the query feature, i.e., the segment similarity, which is used to measure the degree of semantic matching between the corresponding video segment and the text query.

[0026] S140. Perform weighted aggregation on all segment similarities of the same video to be retrieved to obtain the overall similarity between the video to be retrieved and the text query.

[0027] Among them, overall similarity can refer to a quantitative value that comprehensively reflects the semantic similarity between the entire video to be retrieved and the text query. It is the result of fusing the semantic information of all segments of the video and is used for the final ranking of search results.

[0028] In this embodiment of the invention, for any video to be retrieved, the segment similarity corresponding to its video segments can be normalized to obtain the weight corresponding to each video segment. Segments with high similarity will receive a larger weight, while the weight of segments with low similarity will approach zero. Next, the segment similarities and their weights within the same video to be retrieved are weighted and summed to obtain the overall similarity between the corresponding video to be retrieved and the text query.

[0029] S150. Sort each video to be retrieved according to the overall similarity and output the retrieval results.

[0030] In this embodiment of the invention, all videos to be retrieved can be sorted in descending order according to their overall similarity, and the top few videos can be used as the search results. Furthermore, auxiliary information such as the name of each video, shooting time, similarity score, and segment attention weight heatmap can be output, so that users can understand which segments in the video are most relevant to the text query, thereby enhancing the credibility of the search results.

[0031] The technical solution of this invention performs segmented cross-transformation on the text query in a manner consistent with the video side, and utilizes a pre-stored set of video segment features that integrates intra-segment temporal information and cross-segment semantic associations. Combined with temporal alignment matching and weighted aggregation strategies, it enables rapid semantic-level retrieval of power service videos. This solution avoids the cumbersome computation of repeatedly extracting video features online, enhances feature expression through cross-segment semantic associations, and overcomes the problem of inconsistent query and video lengths through flexible temporal alignment. This significantly improves retrieval accuracy and response speed, effectively distinguishes video content of different semantic categories, and overcomes the shortcomings of traditional methods such as semantic fragmentation, insufficient accuracy, and low efficiency.

[0032] Furthermore, based on the above embodiments of the invention, such as Figure 2 As shown, semantic feature extraction and segmented cross-transformation are performed on the acquired text query to obtain query features, including: S210. Obtain the text query input by the user.

[0033] In this embodiment of the invention, natural language query requests submitted by users through search interfaces, API interfaces, etc. can be received. The query content can be a semantic description in the context of power services, such as "transformer arc-extinguishing chamber oil leakage detection" or "GIS equipment gas pressure standard operating procedure".

[0034] S220. Preprocess the text query to obtain a word sequence.

[0035] Among them, a word sequence can refer to a sequence of multiple independent words arranged in order after continuous natural language text has been segmented, and it is the basic input unit of natural language processing.

[0036] In this embodiment of the invention, preprocessing operations such as word segmentation, stop word removal, case unification, and special symbol removal can be performed on the original text query to obtain an ordered word sequence composed of effective words, thus preserving the core semantic information of the text.

[0037] S230. Input the word sequence into the pre-trained language model to obtain the semantic feature vector.

[0038] Among them, the pre-trained language model can refer to a deep learning language model that has been pre-trained on a large-scale general corpus, such as at least BERT_base (basic bidirectional encoder representation model), which is used to encode word sequences into feature vectors rich in semantic information.

[0039] In this embodiment of the invention, the preprocessed word sequence can be input into a pre-trained language model (such as BERT_base). This model can perform context encoding on the word sequence through a multi-layer Transformer structure, outputting a fixed-dimensional semantic feature vector, denoted as... (Query represents the preprocessed sequence of words in a text query.) (Representing the BERT_base model), this vector as a whole expresses the deep semantics of the text query.

[0040] S240. Expand the semantic feature vector according to the preset query time sequence length to obtain the expanded semantic matrix.

[0041] The preset query time series length can refer to the number of rows pre-defined for expanding the text semantic feature vector into a time series matrix. This length can be matched with the time series length range of video segmentation features to match the time series structure of the video segmentation features. The expanded semantic matrix can refer to a two-dimensional time series feature matrix obtained by repeatedly expanding a one-dimensional semantic feature vector. Its number of rows is equal to the preset query time series length, and its number of columns is equal to the dimension of the semantic feature vector.

[0042] In this embodiment of the invention, to address the mismatch between the dimensions of the text semantic feature vector (single vector) and the video segmentation features (temporal matrix), the aforementioned extracted one-dimensional semantic feature vector can be expanded into a two-dimensional extended semantic matrix through repeated copying. The number of rows in this matrix is ​​equal to the preset query time sequence length. The number of columns equals the dimension of the semantic feature vectors. Understandably, each row of the expanded semantic matrix contains the same semantic feature vectors, thus simulating a feature sequence with a temporal dimension.

[0043] S250. Perform a segmented cross transformation on the extended semantic matrix to obtain the query features.

[0044] In this embodiment of the invention, the extended semantic matrix can be... Perform the same segmented cross-transformation operation as the video segmentation features, including steps such as temporal coding, cross-segment correlation information extraction and fusion, and dimensionality compression, to finally obtain query features with dimensions completely consistent with the video segmentation features. The implementation process of segmented crossover will be described in detail below.

[0045] This embodiment achieves semantic alignment between the text modality and the video modality by converting the user-input natural language text query into a temporal feature matrix consistent with the video segmentation feature space. By employing a temporal expansion operation, single-sentence text is given a temporal dimension, facilitating flexible matching using dynamic time warping algorithms and avoiding matching failures caused by inconsistencies between the query and video lengths. Furthermore, the application of segmented cross-transformation endows the text query features with similar cross-segment association capabilities to video features, further enhancing the semantic consistency of the retrieval.

[0046] Furthermore, based on the above embodiments of the invention, the features of each video segment of the video to be retrieved are obtained in the following manner: The video to be retrieved is subjected to frame sampling processing, and the visual features of each frame are extracted to generate a frame feature sequence. Adaptive segmentation of the frame feature sequence yields multiple video segments; Perform segmented cross-transformation on multiple video segments to obtain their respective video segment features.

[0047] Sampling processing refers to the operation of extracting image frames from the original video at fixed intervals (e.g., one frame at a time) to reduce the frame rate and decrease the computational load of subsequent processing. Visual features refer to high-dimensional numerical vectors extracted from video frame images that characterize the essence of the image content. These can be generated using deep learning models such as Convolutional Neural Networks (CNNs) and can include multiple layers of semantic information, such as texture, shape, object category, and spatial location. Frame feature sequences are one-dimensional sequences formed by arranging the visual features of each frame in the video according to the chronological order of playback, fully representing the dynamic changes of the video content over time. Adaptive segmentation refers to the process of automatically dividing the video into segments based on the semantic changes of the video content itself, without pre-setting fixed segment lengths or time intervals, ensuring semantic consistency within each video segment. Video segmentation refers to dividing the original video into several consecutive frame intervals through adaptive segmentation, where the frames within each interval are highly semantically related.

[0048] In this embodiment of the invention, the process of performing offline preprocessing on the video to be detected to extract video segmentation features that can be directly used for subsequent online retrieval specifically includes: S1. Input the video to be searched Perform frame sampling processing, for example, extract one frame every other frame, to obtain a downsampled frame image sequence. (Where i represents the video ID to be retrieved, and M represents the total number of frames after downsampling). Then, for each frame... The input is then fed into a pre-trained convolutional neural network (CNN), and after global average pooling, it outputs a fixed-dimensional visual feature vector. The extraction process is as follows: In the formula, For the j-th frame of the i-th video to be retrieved, including RGB three channels, the pixel values ​​are normalized to [0,1]; CNN is a convolutional neural network; GlobalAvgPool represents the global average pooling operation.

[0049] Finally, the visual features of all frames are arranged in chronological order to form a frame feature sequence. .

[0050] S2. Based on statistical features of the frame feature sequence (such as mean, standard deviation, median, etc.), adaptive segmentation is performed to obtain multiple video segments. (K is the total number of segments), each video segment Corresponding to a continuous frame index range ( And it satisfies the previous index of the first segment start frame. The end frame index of the last segment .

[0051] S3, For each video segment It can extract all the frame features contained therein to form the original segmented feature matrix. Then, the original segmented feature matrix Perform segmented cross-transformation to generate corresponding video segment features. .

[0052] Furthermore, based on the above embodiments of the invention, such as Figure 3 As shown, adaptive segmentation of the frame feature sequence yields multiple video segments, including: S310. Determine the semantic similarity between each pair of adjacent frame features in the frame feature sequence to obtain a similarity sequence.

[0053] In this embodiment of the invention, for the downsampled frame feature sequence The semantic similarity of each pair of adjacent frames can be calculated sequentially. Semantic similarity can be measured using more than just cosine similarity. Furthermore, if a frame's features... If all values ​​are zero (e.g., a black screen frame), skip that frame and replace it with the characteristics of the previous valid frame. This avoids calculation errors when the denominator is zero. Semantic similarity It can be represented as follows: Finally, a similarity order of length M-1 is obtained. .

[0054] S320. Determine the segmentation threshold based on the statistical characteristics of similarity sequences.

[0055] In this embodiment of the invention, statistical features of similarity sequences can be calculated, and the video to be retrieved can be determined based on these statistical features. Dedicated segmentation threshold This threshold can automatically adapt to the semantic changes in the video, eliminating the need for manually setting a fixed threshold. In one specific embodiment, the statistical feature may include the mean. and standard deviation : Next, based on the 3σ criterion (Raida criterion), the corresponding segmented thresholds are set. This means that when the similarity between adjacent frames is more than three standard deviations below the mean, the semantics have changed significantly and the frame should be considered as a candidate position for segmentation.

[0056] S330. Mark the next frame in the similarity sequence whose value is less than the segmentation threshold as the initial segmentation point.

[0057] In this embodiment of the invention, the entire similarity sequence can be traversed, and the similarity value at each position can be... With segmented threshold Compare the similarity values ​​at a certain location. Less than When the semantic difference between two adjacent frames at a given position is significant, it indicates that there is a semantic jump between frame j and frame j+1, and frame j+1 can be marked as an initial segmentation point. After traversal, multiple initial segmentation points can be obtained.

[0058] S340. Process the initial segmentation points according to the preset segmentation rules to obtain multiple video segments.

[0059] Among them, the preset segmentation rules can refer to the pre-set rules for optimizing the initial segmentation points. Specifically, they can include consecutive segmentation point filtering rules and short segment merging rules. For example, if a preset number or more initial segmentation points appear consecutively, only the first initial segmentation point is retained; if the number of frames contained in the divided video segment is less than the preset minimum number of frames, the video segment is merged into the adjacent previous video segment.

[0060] In this embodiment of the invention, all the initial segmentation points obtained above can be optimized according to a preset segmentation rule. Specifically, if a preset number (e.g., 3) or more initial segmentation points appear consecutively (e.g., in a fast video switching scene), only the first initial segmentation point is retained, and subsequent consecutive segmentation points are deleted to avoid over-segmentation. If, after dividing the video according to the retained segmentation points, the number of frames in the resulting video segments is less than a preset minimum number of frames (e.g., 5 frames), the video segment is merged into the adjacent preceding video segment to ensure that each video segment contains enough frames to carry complete semantics. Finally, the video to be retrieved is obtained. Corresponding adaptive segmentation results In one embodiment, each video segment It can contain the number of frames (i.e., the sequence length). To balance semantic integrity and computational efficiency.

[0061] This embodiment, through the aforementioned adaptive segmentation scheme, dynamically determines segment positions based on semantic changes in video content, avoiding semantic fragmentation caused by fixed time window division. By employing statistical features based on similarity sequences to determine segmentation thresholds, it exhibits adaptability to different video content—videos with drastic content changes generate larger standard deviations and higher thresholds, thereby reducing meaningless segmentation; videos with smooth content generate lower thresholds, ensuring sensitivity to semantic changes. Furthermore, the continuous segment point retention rule effectively addresses false detections caused by rapidly switching scenes (such as flickering and transitions), while the short segment merging rule ensures that each segment possesses sufficient semantic information, facilitating subsequent temporal coding and cross-segment association extraction. Ultimately, it provides high-quality segments with semantic uniformity and appropriate length for subsequent segmented cross-transformation, reducing the complexity of long video processing and improving the semantic accuracy of retrieval.

[0062] Furthermore, based on the above embodiments of the invention, such as Figure 4 As shown, a segmented cross-transform is performed on the video segments to obtain the corresponding video segment features, including: S410. Perform temporal encoding on the frame feature sequence contained in the video segment to obtain temporal enhanced features. The temporal enhanced features integrate the contextual temporal dependencies within the video segment.

[0063] In this embodiment of the invention, a bidirectional Long Short-Term Memory (LSTM) network can be used for video segmentation. The original piecewise feature matrix Temporal encoding is performed, and the forward LSTM in this bidirectional LSTM (BiLSTM) processes the frame sequence in chronological order, outputting the forward hidden state at each time step t. The backward LSTM processes the frame sequence in reverse chronological order and outputs the backward hidden state. By concatenating the hidden states in both directions, the temporal augmentation feature vector for each frame can be obtained. Then, the temporal augmentation feature vectors of all frames are arranged in order to form the original segmented feature matrix. Corresponding temporal enhancement features (matrix) , means as follows: In the formula, , This indicates a splicing operation.

[0064] S420. Determine the first cross-covariance matrix between the video segment and the temporal enhancement feature corresponding to the previous video segment, and the second cross-covariance matrix between the video segment and the temporal enhancement feature corresponding to the next video segment; wherein, if the video segment is the first segment, the first cross-covariance matrix is ​​a zero matrix; if the video segment is the last segment, the second cross-covariance matrix is ​​a zero matrix.

[0065] In this embodiment of the invention, for each video segment The cross-covariance matrix between the current video segment and its temporal enhancement features can be calculated to quantify the cross-segment semantic association strength. Specifically, for non-boundary segments (2≤k≤K-1), the cross-covariance matrix between the current video segment and its temporal enhancement features can be calculated. Segmented from the previous video The first cross covariance matrix between and the current video segmentation Segmented with the next video The second cross covariance matrix between For the first segment (k=1), since there is no preceding segmentation, the first cross-covariance matrix is... Set as an all-zero matrix, and calculate only the second cross covariance matrix. For the final segment (k=K), since there is no subsequent segmentation, the second cross-covariance matrix is... Set as an all-zero matrix, and calculate only the first cross covariance matrix. .

[0066] S430. Flatten the first cross-covariance matrix and the second cross-covariance matrix into one-dimensional vectors and then concatenate them to obtain a one-dimensional correlation feature vector.

[0067] In this embodiment of the invention, the first cross covariance matrix can be... and Converting two-dimensional matrices to one-dimensional vectors respectively and During the transformation, the total number and relative order of matrix elements must be kept unchanged. Then, the two one-dimensional vectors are concatenated end-to-end to obtain a one-dimensional correlation feature vector that simultaneously contains forward and backward cross-segment correlation information. This vector represents the global correlation between the current segment and the segments before and after it across all feature dimensions.

[0068] S440. Concatenate each frame feature vector in the temporal enhancement feature with a one-dimensional associated feature vector to obtain the fused feature.

[0069] In this embodiment of the invention, for temporal enhancement features Each frame feature vector (That is, each row of the matrix) can be associated with the one-dimensional feature vector obtained in step S430. Dimensional concatenation is performed to obtain a high-dimensional fusion feature that simultaneously contains single-frame temporal information and global cross-segment correlation information. .

[0070] S450. Perform dimensional compression on the fused features to obtain video segmentation features with the same dimensions as the visual features.

[0071] In this embodiment of the invention, high-dimensional features of each frame can be fused using a multi-layer neural network. Dimensional compression is performed, and the final output is a single-frame video compressed feature with the same dimensions as the original single-frame visual features. In one embodiment, a two-layer fully connected network can be used to fuse high-dimensional features. Dimensional compression is performed as follows: In the formula, and These are the weight matrix and bias vector of the first fully connected layer, respectively. and These are the weight matrix and bias vector of the second fully connected layer, respectively; ReLU is the activation function to prevent gradient vanishing.

[0072] Finally, the compressed features of all frames are arranged in order to obtain the video segmentation feature matrix. .

[0073] This embodiment, through the aforementioned segmented cross-transformation process, organically integrates the local frame features of a single video segment with the global semantic information of preceding and following segments, significantly enhancing the semantic expressive power of video segment features. Specifically, through bidirectional LSTM temporal coding, the representation of each frame contains the contextual dependencies within the segment, solving the problem of insufficient information in a single frame. The introduction of a cross-segment covariance matrix enables the current segment to perceive the overall feature correlation between preceding and following segments, thereby integrating semantic cues scattered across different segments into the same feature space, avoiding the semantic fragmentation caused by traditional fixed segmentation. Zero matrix processing of the first and last segments ensures the uniformity and robustness of the algorithm. Flattening and splicing operations compress high-dimensional covariance information into learnable features, followed by dimensionality compression through a fully connected network, reducing computational overhead while retaining key correlation information. The resulting video segment features possess rich intra-segment temporal information and integrate cross-segment semantic correlations, providing a high-quality semantic foundation for subsequent dynamic time warping matching and accurate retrieval, and serving as the core support for achieving second-level, high-precision power service video retrieval.

[0074] It's important to understand that the process of obtaining query features by performing a piecewise cross-transformation on the extended semantic matrix of a text query is exactly the same as the processing on the video side. Specifically, firstly, the extended semantic matrix... The input is a bidirectional LSTM for temporal encoding to obtain temporal enhancement features. Then, the cross-covariance matrices with the preceding and following segments are calculated separately. However, since the text query does not have actual preceding and following segments, the first and second cross-covariance matrices are set to zero matrices. These two zero matrices are then flattened and concatenated into a one-dimensional correlation feature vector. This vector is then concatenated with the feature vectors of each frame in the temporal enhancement features to obtain the fused features. Finally, a two-layer fully connected network compresses the fused features to the same dimension as the video segmentation features, ultimately outputting the query features. This design ensures consistency and comparability between text queries and video segments in the feature space.

[0075] Furthermore, based on the above embodiments of the invention, for each video segment feature of each video to be retrieved, a temporal alignment matching algorithm is used to determine the alignment distance between the video segment feature and the query feature, and the segment similarity between the video segment feature and the query feature is determined based on the alignment distance, including: A dynamic time warping algorithm is used to determine the alignment distance between video segment features and query features; The alignment distance is normalized to obtain the initial similarity, and the initial similarity is corrected according to the temporal length difference between the video segment features and the query features to obtain the segment similarity.

[0076] In this embodiment of the invention, the temporal alignment matching algorithm may include a dynamic time warping (DTW) algorithm. The process of determining the alignment distance between video segment features and query features based on the DTW algorithm, and then determining the segment similarity between the two, specifically includes: S1. For the current video to be searched Feature matrix of each video segment and query features First, calculate the Euclidean distance between all frame pairs between them to obtain the frame-level distance matrix D, which has a size of Each element in the matrix This represents the distance between the features of frame t1 in the video segment and the features of frame t2 in the query features. The smaller the distance value, the higher the semantic similarity between the two frames. The segmented frame index is used for this purpose. Query frame index L is the dimension of the feature vector, that is, the total number of feature components contained in each frame feature vector or query feature vector.

[0077] Then, the cumulative distance matrix E is calculated using dynamic programming, with the recursive formula as follows: Initial conditions are The accumulated distance in the first row is The accumulated distance in the first column is .

[0078] Finally, the bottom right element of the cumulative distance matrix E This is the alignment distance obtained by the DTW algorithm, denoted as... The smaller the distance value, the smaller the overall difference between the two feature sequences under optimal alignment.

[0079] S2. The alignment distance is calculated using the following normalization formula. Convert to initial similarity : In the formula, is a pre-defined normalization coefficient used to ensure that the similarity value falls within the [0,1] interval; It is a very small positive number, used to avoid the denominator being zero.

[0080] Next, considering the video segmentation feature matrix and query features Differences in temporal length can also affect the rationality of matching; sequences with significantly different lengths may not be a semantic match even if the DTW distance is small. Therefore, a length penalty factor is introduced. : Finally, the segment similarity is obtained as follows: This similarity score comprehensively reflects the quality of feature alignment and the consistency of temporal length; the larger the value, the more relevant the current video segment is to the text query.

[0081] This embodiment employs a dynamic time warping algorithm to effectively address the inconsistency in temporal length between video segmentation features and text query features. By using non-linear alignment to find the optimal matching path, it avoids matching failures caused by timeline misalignment, significantly improving retrieval robustness. Simultaneously, normalization converts the alignment distance into an intuitive similarity value, facilitating subsequent comparison and aggregation. The introduction of a temporal length difference correction term penalizes sequences with excessively large length differences, preventing false high similarity due to extreme length mismatches and further enhancing retrieval accuracy.

[0082] Furthermore, based on the above embodiments of the invention, the weighted aggregation of the segment similarities of the same video to be retrieved is performed to obtain the overall similarity between the video to be retrieved and the text query, including: Normalization is performed on the similarity of all segments of the same video to be retrieved to obtain the segment weights corresponding to each video segment. The segment similarity and segment weight of each video segment are weighted and summed to obtain the overall similarity between the video to be retrieved and the text query.

[0083] In this embodiment of the invention, the process of determining the overall similarity between the video to be retrieved and the text query through weighted aggregation specifically includes: S1. For the video to be searched Assume there are K video segments in total, and the segment similarity between each segment and the text query Q has been calculated in the previous stage. Then, the softmax normalization function can be used to normalize them, so as to convert the segment similarity values ​​into corresponding segment weights: After normalization, video segments with higher similarity will receive greater weight, while video segments with lower similarity will receive less weight.

[0084] S2. Based on the segment similarity corresponding to each video segment With segmented weights The following formula can be used to determine the video to be retrieved. Overall similarity with text query Q : This video-level similarity score integrates the semantic matching of all segments within the video, and the weighting method can highlight the segments most relevant to the query while weakening the influence of irrelevant or low-relevance segments.

[0085] This embodiment, through the aforementioned weighted aggregation process, effectively merges the similarities of multiple segments within a video into a single video-level similarity score, avoiding information loss or unfairness issues arising from simple averaging or taking the maximum value. Through softmax normalization, segments with higher similarity receive significantly higher weights, ensuring that the overall video similarity is primarily determined by the segments most relevant to the query. This aligns with the practical needs of video retrieval—a video containing segments highly matching the query should be considered relevant. Furthermore, the exponential operation of softmax amplifies the advantages of high-segment similarities and suppresses interference from low-segment similarities, improving the accuracy and robustness of the retrieval.

[0086] Furthermore, it can be based on overall similarity. For all videos to be searched Sort the videos in descending order and return the top few videos as output. Also output key information for each video, such as video name, shooting time, segmented attention weight heatmap, similarity score, etc.

[0087] Furthermore, based on the above embodiments, the technical solution of this embodiment also includes: When a data update is detected in the preset power service video library, the changed video data is extracted; If the changed video data is a new video, then the new video segment feature set corresponding to the new video is determined, and the new video segment feature set is synchronized to the preset semantic feature database; If the semantic tags of the video data are changed to those of the existing video, the changed semantic tags will be synchronized to the preset semantic feature database.

[0088] In this embodiment of the invention, the following data update process is also included: S1 can detect changes in the status of a preset power service video library in real time or periodically. The detection content may include: whether new video files have been added to the library, whether existing video files have been deleted, and whether the semantic tags of existing videos (such as equipment type, fault description and other metadata) have been modified manually or automatically.

[0089] S2. Once a data update event is detected, the system immediately extracts the specific details of the change. Specifically, for newly added videos, the video file itself and its storage path can be extracted; for semantic tag changes, the unique identifier of the corresponding video and the changed tag information can be extracted.

[0090] S3. If the video data is changed to a new video, the method described in the above embodiments is used to determine the new video segment feature set corresponding to the new video and store it in the preset semantic feature database, so that the video becomes a searchable object.

[0091] S4. If the video data is changed to the semantic tags of existing videos, the corresponding record for the video in the preset semantic feature database is located directly, and its tag field is updated. It is important to understand that this process does not require recalculating the segmentation features of the video, because the segmentation features only depend on the visual content of the video and are unrelated to the manually labeled semantic tags.

[0092] Furthermore, after completing the video retrieval, the system can leverage the dedicated transmission channel for power service video semantic retrieval to encapsulate incremental data containing key semantic feature parameters of the video and retrieval matching results in a unified semantic retrieval data format, enabling real-time interaction with the power service video cloud retrieval center. The system can continuously monitor the dynamic changes in the power service video database. When new video data or updates to existing video semantic tags are detected, the incremental acquisition module is triggered to accurately extract the video semantic change data. This data is then uploaded to the distributed power service video semantic database cluster via an encrypted transmission channel. Combined with the real-time semantic retrieval inference requirements of the segmented cross-transformation model, the system updates semantic feature data and persistently stores retrieval matching results. This provides dynamic and accurate semantic data support for the dynamic tracking and analysis of power service video semantics, the real-time iteration of parameters in the segmented cross-transformation model, and online decision-making in power service video semantic retrieval.

[0093] To facilitate understanding of this solution by those skilled in the art, this embodiment selects five segments each of three typical power service videos (equipment inspection, fault repair, and operation training). Each video segment is downsampled to 750 frames. Using the text query "transformer fault detection" as the retrieval target, data that closely matches the distribution of real features is simulated, and four types of key graphic results are output to intuitively verify the retrieval effect of the model.

[0094] like Figure 5As shown in the figure, this diagram visually presents the changes in inter-frame semantic similarity and the adaptive segmentation effect of the transformer fault repair video. The horizontal axis represents the inter-frame pairing index (frame j and frame j+1), and the vertical axis represents the inter-frame cosine similarity (the closer the value is to 1, the more semantically coherent the video). The red dashed line represents the segmentation threshold θ, calculated based on the 3σ criterion and uniformly set to 0.68 in this embodiment. The red solid dots represent the segmentation points identified by the system (finally determined to be frames 201, 451, and 601). It can be clearly seen from the figure that the segmentation points all correspond to the regions where the inter-frame similarity decreases significantly, which precisely match the semantic segmentation of "fault discovery - fault detection - fault recording" in the video. Ultimately, the 750 frames of video are divided into 4 semantically unified segments (each segment containing 200, 250, 150, and 150 frames respectively), which avoids the semantic fragmentation caused by fixed segmentation and reduces the computational complexity for subsequent feature extraction.

[0095] like Figure 6 As shown, this figure is a heatmap of the cross-covariance matrix of the second segment (the core stage of the fault) and the segments before and after it (visualizing the first 50 dimensions of features). This represents the first cross-covariance matrix between the current segment and the previous segment. This represents the second cross-covariance matrix between the current segment and the next segment. In the heatmap, the color intensity corresponds to the magnitude of the covariance value; red indicates high covariance (strong feature correlation), and blue indicates low covariance (weak feature correlation). The left image shows the first cross-covariance matrix between the current segment and the previous segment (fault preparation stage), and the right image shows the second cross-covariance matrix between the current segment and the next segment (fault recording stage). It can be observed that the right image has significantly more red high-covariance areas than the left image, indicating a stronger semantic feature correlation between the core fault stage and the fault recording stage (such as semantic continuity in fault phenomenon description and data recording). This verifies that cross-segment cross-transformation can effectively capture the global semantic correlation between segments, solving the problem of one-sided semantics in a single segment, and providing more comprehensive feature support for subsequent accurate retrieval.

[0096] like Figure 7 and 8 As shown, the two figures visually verify the semantic matching effect of video segmentation and text query. Specifically, Figure 7 The DTW matching distance matrix and optimal path map are shown below: The grayscale heatmap represents the frame-level Euclidean distance between the segmented features (250 frames) and the query features (20 frames) (the darker the color, the more similar the frame-level semantics). The white curve is the optimal temporal alignment path found by the DTW algorithm, which successfully achieves flexible matching between long and short sequences. The final calculated unnormalized cumulative DTW distance is 142.36. Figure 8The segmented similarity histogram shows that after normalizing the DTW distance and correcting it with a length penalty, the semantic similarity between this segment and the query "transformer fault detection" is 0.6189 (close to 0.7, which is in the high similarity range). This clearly indicates that the segment contains semantic information that is highly relevant to the query (transformer fault detection scenario), verifying that DTW matching can effectively handle the problem of time series length differences. At the same time, the features extracted by segmented cross-transformation have strong semantic expressive power.

[0097] like Figure 9 As shown in the figure, this is the search ranking result of 15 power service videos. The horizontal axis represents the video index and category labels sorted in descending order of similarity, and the vertical axis represents the final semantic similarity between the video and the query "transformer fault detection". It can be clearly observed from the figure that: the top 5 ranked videos are all fault-related videos (red bars), with similarity values ​​between 0.831 and 0.915. Among them, the video "Fault 2" has the highest similarity (0.915), corresponding to a video containing a complete transformer fault detection process in a real-world scenario; videos 6-10 are inspection-related videos (blue bars), with similarity values ​​between 0.285 and 0.356 (because inspection videos may include transformer visual inspections, showing a weak correlation with the query); videos 11-15 are training-related videos (green bars), with similarity values ​​only between 0.176 and 0.213 (having very little semantic correlation with fault detection). The results fully validate the model's retrieval accuracy—it can accurately identify and rank video categories that are highly relevant to the query semantics, while effectively distinguishing between weakly related and unrelated videos. The retrieval recall rate (the top 5 are all relevant videos) reaches 100%, and the precision rate reaches 100%, fully meeting the accurate retrieval needs in the power service scenario.

[0098] Figure 10 This is a schematic diagram of the structure of a power service video retrieval device provided in an embodiment of the present invention. Figure 10 As shown, the device includes: The query feature determination module 51 is used to extract semantic features and perform segmented cross-transformation on the acquired text query to obtain query features; The video segmentation feature acquisition module 52 is used to read the video segmentation feature set corresponding to each video to be retrieved from the preset semantic feature database. The matching module 53 is used to determine the alignment distance between the video segment features and the query features for each video segment feature of each video to be retrieved using a temporal alignment matching algorithm, and to determine the segment similarity between the video segment features and the query features based on the alignment distance. The overall similarity determination module 54 is used to perform weighted aggregation of the segment similarities of the same video to be retrieved, so as to obtain the overall similarity between the video to be retrieved and the text query. The search results output module 55 is used to sort each video to be searched according to the overall similarity and output the search results.

[0099] Furthermore, based on the above embodiments of the invention, the query feature determination module 51 is specifically used for: Get the text query input by the user; Preprocess the text query to obtain a word sequence; Input the word sequence into the pre-trained language model to obtain semantic feature vectors; The semantic feature vector is expanded according to the preset query time sequence length to obtain the expanded semantic matrix; The extended semantic matrix is ​​subjected to a segmented cross-transformation to obtain the query features.

[0100] Furthermore, based on the above embodiments of the invention, the features of each video segment of the video to be retrieved are obtained in the following manner: The video to be retrieved is subjected to frame sampling processing, and the visual features of each frame are extracted to generate a frame feature sequence. Adaptive segmentation of the frame feature sequence yields multiple video segments; Perform segmented cross-transformation on multiple video segments to obtain their respective video segment features.

[0101] Furthermore, based on the above embodiments of the invention, adaptive segmentation is performed on the frame feature sequence to obtain multiple video segments, including: Determine the semantic similarity between every two adjacent frame features in the frame feature sequence to obtain a similarity sequence; Segmentation thresholds are determined based on statistical characteristics of similarity sequences; Mark the next frame in the similarity sequence whose value is less than the segmentation threshold as the initial segmentation point; The initial segmentation points are processed according to the preset segmentation rules to obtain multiple video segments; The preset segmentation rules include: If a preset number or more initial segmentation points appear consecutively, only the first initial segmentation point will be retained. If the number of frames in a video segment is less than the preset minimum number of frames, the video segment will be merged into the adjacent previous video segment.

[0102] Furthermore, based on the above embodiments of the invention, a segmented cross-transform is performed on the video segments to obtain the corresponding video segment features, including: Temporal encoding is performed on the frame feature sequences contained in the video segments to obtain temporal enhanced features, which incorporate the contextual temporal dependencies within the video segments; The first cross-covariance matrix between the video segment and the temporal enhancement feature corresponding to the previous video segment, and the second cross-covariance matrix between the video segment and the temporal enhancement feature corresponding to the next video segment are determined respectively; wherein, if the video segment is the first segment, the first cross-covariance matrix is ​​a zero matrix; if the video segment is the last segment, the second cross-covariance matrix is ​​a zero matrix. The first cross-covariance matrix and the second cross-covariance matrix are flattened into one-dimensional vectors and then concatenated to obtain a one-dimensional associated feature vector. Each frame feature vector in the temporal enhancement features is concatenated with a one-dimensional associated feature vector to obtain the fused features; The fusion features are subjected to dimensionality compression to obtain video segmentation features with the same dimensionality as the visual features.

[0103] Furthermore, based on the above embodiments of the invention, the timing alignment matching algorithm includes a dynamic time warping algorithm; the matching module 53 is specifically used for: A dynamic time warping algorithm is used to determine the alignment distance between video segment features and query features; The alignment distance is normalized to obtain the initial similarity, and the initial similarity is corrected according to the temporal length difference between the video segment features and the query features to obtain the segment similarity.

[0104] Furthermore, based on the above embodiments of the invention, the overall similarity determination module 54 is specifically used for: Normalization is performed on the similarity of all segments of the same video to be retrieved to obtain the segment weights corresponding to each video segment. The segment similarity and segment weight of each video segment are weighted and summed to obtain the overall similarity between the video to be retrieved and the text query.

[0105] Furthermore, based on the above embodiments of the invention, the power service video retrieval device further includes: The detection module is used to extract the changed video data when a data update is detected in the preset power service video library; The first processing module is used to determine the new video segment feature set corresponding to the new video if the changed video data is a new video, and to synchronize the new video segment feature set to the preset semantic feature database. The second processing module is used to synchronize the changed semantic tags to the preset semantic feature database if the video data is changed to the semantic tags of the existing video.

[0106] The power service video retrieval device provided in this embodiment of the invention can execute the power service video retrieval method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0107] Figure 11 A schematic diagram of an electronic device 60 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0108] like Figure 11 As shown, the electronic device 60 includes at least one processor 61 and a memory, such as a read-only memory (ROM) 62 and a random access memory (RAM) 63, communicatively connected to the at least one processor 61. The memory stores computer programs executable by the at least one processor. The processor 61 can perform various appropriate actions and processes based on the computer program stored in the ROM 62 or loaded from storage unit 68 into the RAM 63. The RAM 63 can also store various programs and data required for the operation of the electronic device 60. The processor 61, ROM 62, and RAM 63 are interconnected via a bus 64. An input / output (I / O) interface 65 is also connected to the bus 64.

[0109] Multiple components in electronic device 60 are connected to I / O interface 65, including: input unit 66, such as keyboard, mouse, etc.; output unit 67, such as various types of monitors, speakers, etc.; storage unit 68, such as disk, optical disk, etc.; and communication unit 69, such as network card, modem, wireless transceiver, etc. Communication unit 69 allows electronic device 60 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0110] Processor 61 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 61 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 61 performs the various methods and processes described above, such as the power service video retrieval method.

[0111] In some embodiments, the power service video retrieval method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 68. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 60 via ROM 62 and / or communication unit 69. When the computer program is loaded into RAM 63 and executed by processor 61, one or more steps of the power service video retrieval method described above may be performed. Alternatively, in other embodiments, processor 61 may be configured to perform the power service video retrieval method by any other suitable means (e.g., by means of firmware).

[0112] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0113] In some embodiments, the power service video retrieval method may be implemented as a computer program, which is implicitly included in a computer program product. When executed by a processor, the computer program implements the power service video retrieval method of the present invention. The computer program product can be understood as a software product that primarily implements its solution through a computer program. The computer program used to implement the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer program causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer program may be executed entirely on a machine, partially on a machine, partially on a remote machine as a standalone software package, or entirely on a remote machine or server.

[0114] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0115] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0116] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0117] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0118] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0119] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for retrieving video feeds related to electricity services, characterized in that, The method includes: Semantic feature extraction and segmented cross-transformation are performed on the acquired text query to obtain query features; Read the video segment feature set corresponding to each video to be retrieved from the preset semantic feature database; For each video segment feature of each of the videos to be retrieved, a temporal alignment matching algorithm is used to determine the alignment distance between the video segment feature and the query feature, and the segment similarity between the video segment feature and the query feature is determined based on the alignment distance; The weighted aggregation of all segment similarities of the same video to be retrieved yields the overall similarity between the video to be retrieved and the text query. The videos to be retrieved are sorted according to the overall similarity, and the retrieval results are output.

2. The method according to claim 1, characterized in that, The semantic feature extraction and segmented cross-transformation of the acquired text query to obtain query features include: Obtain the text query input by the user; The text query is preprocessed to obtain a word sequence; The word sequence is input into a pre-trained language model to obtain semantic feature vectors; The semantic feature vector is expanded according to a preset query time sequence length to obtain an expanded semantic matrix; The segmented cross transformation is performed on the extended semantic matrix to obtain the query features.

3. The method according to claim 1, characterized in that, The features of each video segment of the video to be retrieved are obtained in the following way: The video to be retrieved is subjected to frame sampling processing, and the visual features of each frame are extracted to generate a frame feature sequence. The frame feature sequence is adaptively segmented to obtain multiple video segments; The segmented cross-transform is performed on each of the multiple video segments to obtain their respective video segment features.

4. The method according to claim 3, characterized in that, The adaptive segmentation of the frame feature sequence to obtain multiple video segments includes: Determine the semantic similarity between every two adjacent frame features in the frame feature sequence to obtain a similarity sequence; The segmentation threshold is determined based on the statistical characteristics of the similarity sequence; Mark the next frame in the similarity sequence whose value is less than the segmentation threshold as the initial segmentation point; The initial segmentation points are processed according to preset segmentation rules to obtain multiple video segments; The preset segmentation rules include: If a preset number or more initial segmentation points appear consecutively, only the first initial segmentation point will be retained. If the number of frames in a video segment is less than the preset minimum number of frames, the video segment will be merged into the adjacent preceding video segment.

5. The method according to claim 3, characterized in that, Performing the segmented cross-transform on the video segments to obtain the corresponding video segment features includes: Temporal encoding is performed on the frame feature sequence contained in the video segment to obtain temporal enhanced features, which incorporate the contextual temporal dependencies within the video segment; The first cross-covariance matrix between the video segment and the temporal enhancement feature corresponding to the previous video segment, and the second cross-covariance matrix between the video segment and the temporal enhancement feature corresponding to the next video segment are determined respectively; wherein, if the video segment is the first segment, the first cross-covariance matrix is ​​a zero matrix; if the video segment is the last segment, the second cross-covariance matrix is ​​a zero matrix. The first cross-covariance matrix and the second cross-covariance matrix are flattened into one-dimensional vectors and then concatenated to obtain a one-dimensional associated feature vector. Each frame feature vector in the temporal enhancement features is concatenated with the one-dimensional associated feature vector to obtain the fused features; The fusion features are subjected to dimensional compression to obtain the video segmentation features with the same dimensions as the visual features.

6. The method according to claim 1, characterized in that, The temporal alignment matching algorithm includes a dynamic time warping algorithm; for each video segment feature of each of the videos to be retrieved, the temporal alignment matching algorithm is used to determine the alignment distance between the video segment feature and the query feature, and the segment similarity between the video segment feature and the query feature is determined based on the alignment distance, including: The alignment distance between the video segmentation features and the query features is determined using the dynamic time warping algorithm. The alignment distance is normalized to obtain an initial similarity, and the initial similarity is corrected according to the temporal length difference between the video segment features and the query features to obtain the segment similarity.

7. The method according to claim 1, characterized in that, The step of weighted aggregation of all segment similarities of the same video to be retrieved to obtain the overall similarity between the video to be retrieved and the text query includes: Normalization is performed on all the segment similarities of the same video to be retrieved to obtain the segment weights corresponding to each video segment; The segment similarity and segment weight corresponding to each video segment are weighted and summed to obtain the overall similarity between the video to be retrieved and the text query.

8. The method according to claim 1, characterized in that, Also includes: When a data update is detected in the preset power service video library, the changed video data is extracted; If the changed video data is a newly added video, then the newly added video segment feature set corresponding to the newly added video is determined, and the newly added video segment feature set is synchronized to the preset semantic feature database; If the changed video data is a semantic tag of an existing video, then the changed semantic tag will be synchronized to the preset semantic feature database.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the power service video retrieval method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the power service video retrieval method according to any one of claims 1-8.