Multi-modal intelligent analysis method and system for power talent capability assessment
By performing time-series alignment and format regularization of multimodal data, weighted fusion, and high-dimensional mapping calculation on the power talent competency assessment method, and combining geometric hashing algorithm for feature difference matching, a personalized scoring rule set is constructed. This solves the problem of insufficient multimodal data processing in existing technologies and achieves a more accurate and suitable comprehensive competency assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAN DIANKE EDUCATION TECH CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-08
AI Technical Summary
Existing power industry talent competency assessment methods suffer from insufficient spatiotemporal correlation and semantic complementarity in multimodal data processing, lack dynamic feature extraction capabilities, resulting in limited accuracy of assessment results. Furthermore, they lack end-to-end deep feature learning mechanisms, making them unable to adapt to complex and ever-changing operating environments.
By dynamically allocating modal weights, the multimodal input data is time-series aligned and formatted, and weighted fusion is performed to generate a multimodal data matrix. A pre-trained neural network is used for high-dimensional mapping calculation, and a geometric hash algorithm is combined to perform feature difference matching and virtual type template mapping. A personalized scoring adjustment rule set is constructed, and finally, multi-dimensional weighted aggregation is performed to generate a comprehensive capability assessment result.
It achieves comprehensive association and feature representation of multimodal data, improves the accuracy and relevance of evaluation results, adapts to the current scoring range of virtual imaging objects, and generates a more realistic comprehensive capability assessment.
Smart Images

Figure CN121998502A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power industry technology, and in particular to a multimodal intelligent analysis method and system for assessing the capabilities of power personnel. Background Technology
[0002] As the scale of power systems continues to expand and technological complexity continues to increase, the demand for skilled professionals in the power industry is becoming increasingly urgent. Traditional methods for assessing talent competence mainly rely on written examinations, expert observation and evaluation, or single-dimensional practical records. These methods have significant shortcomings in terms of objectivity, comprehensiveness, and efficiency. In particular, at the data processing level, existing technologies are unable to meet the requirements of multimodal, high-dimensional data analysis, thus limiting the accuracy of assessment results.
[0003] For example, existing technologies typically employ simple data splicing or weighted averaging methods, lacking effective alignment and deep fusion mechanisms for heterogeneous data such as video, audio, and motion sensing. This coarse fusion method cannot accurately capture the spatiotemporal correlation and semantic complementarity between multimodal data, resulting in the loss of important feature information.
[0004] Secondly, power operations have strong temporal characteristics. Existing methods often use static feature extraction, ignoring the dynamic changes in the operation process and lacking the ability to align time-series data and extract dynamic features. This makes it impossible to accurately model the evolution process and time dependencies of operation behavior.
[0005] Third, existing technologies mostly employ manually designed feature extraction methods, which are often limited to surface phenomena and fail to uncover deeper semantic information from the data. The lack of an end-to-end deep feature learning mechanism results in limited feature representation capabilities, making it unable to adapt to complex and ever-changing operating environments.
[0006] In recent years, although some studies have attempted to apply machine learning methods to the field of talent assessment, these methods also have some limitations: on the one hand, most methods are based on single-modal data and cannot make full use of the complementary advantages of multimodal data; on the other hand, even if multimodal data is used, there is a lack of systematic data processing procedures, especially in key links such as spatiotemporal alignment, feature fusion and personalized calibration. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a multimodal intelligent analysis method and system for assessing the capabilities of power industry personnel. By dynamically allocating modal weights, optimizing data processing, calibrating group differences, and correcting biases, the method improves the matching degree of the assessment.
[0008] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a multimodal intelligent analysis method for assessing the capabilities of power industry personnel, the method comprising: Based on a preset weight configuration strategy, the multimodal input data that has been time-aligned and formatted is weighted and fused to generate a fused multimodal data matrix; based on this matrix, multimodal feature vectors are extracted to characterize operational behavior patterns, interactive response characteristics, and state judgment indicators. The multimodal feature vectors are input into a pre-trained neural network regression model for high-dimensional mapping calculation to obtain an initial capability assessment score. Based on the initial ability assessment score, typical modal features are extracted from speech and action modalities, and the difference measure between them and the baseline patterns in the virtual talent template library is calculated to generate a feature difference vector. The geometric hash algorithm is used to quickly match and map the feature difference vector to obtain a suitable virtual type template. Based on the virtual type template, dynamic region division and merging are performed in the three-dimensional feature space composed of operation behavior patterns, state judgment indicators and interactive response characteristics to generate a scoring interval adapted to the current virtual imaging object; through interval mapping and normalization operations, a personalized scoring adjustment rule set is constructed. The initial ability assessment score is calibrated based on the personalized scoring adjustment rule set to obtain an optimized ability score. The optimized capability scores are weighted and aggregated in multiple dimensions to generate a comprehensive capability assessment result, which covers four virtual imaging dimensions: operational behavior patterns, interactive response characteristics, state judgment indicators, and execution stability.
[0009] Secondly, a multimodal intelligent analysis system for assessing the capabilities of power industry personnel includes: The extraction module is used to perform weighted fusion on the multimodal input data that has been time-aligned and formatted according to a preset weight configuration strategy, and generate a fused multimodal data matrix; based on this matrix, multimodal feature vectors are extracted to represent operation behavior patterns, interaction response characteristics and state judgment indicators; The mapping module is used to input the multimodal feature vectors into a pre-trained neural network regression model for high-dimensional mapping calculation to obtain an initial capability assessment score. The matching module is used to extract typical modal features from speech and action modalities based on the initial ability assessment score, calculate the difference measure between them and the benchmark patterns in the virtual talent template library, and generate a feature difference vector; the feature difference vector is quickly matched and mapped using a geometric hash algorithm to obtain a suitable virtual type template. The rule set module is used to dynamically divide and merge regions in a three-dimensional feature space composed of operation behavior patterns, state judgment indicators and interaction response characteristics according to the virtual type template, and generate a scoring interval adapted to the current virtual imaging object; and to construct a personalized scoring adjustment rule set through interval mapping and normalization operations. The calibration module is used to calibrate the initial ability assessment score based on the personalized scoring adjustment rule set to obtain an optimized ability score; The evaluation results module is used to perform multi-dimensional weighted aggregation of the optimized capability scores to generate comprehensive capability evaluation results. The evaluation results cover four virtual imaging dimensions: operational behavior patterns, interactive response characteristics, state judgment indicators, and execution stability.
[0010] Thirdly, a computing device includes: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.
[0011] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0012] The above-described solution of the present invention has at least the following beneficial effects: By weighted fusion of time-aligned multimodal data to generate matrices and extract features, multimodal information can be integrated to cover key evaluation dimensions, providing a comprehensive and correlated data foundation for subsequent analysis. Multimodal feature vectors are input into a pre-trained neural network for high-dimensional mapping, making the initial capability assessment score more closely aligned with feature representations and adapting to evaluation needs. Typical features are extracted from speech and action modalities, and differences are calculated. Combined with a geometric hash algorithm for rapid matching, this accurately captures feature differences while efficiently finding suitable virtual type templates. Dynamic division and merging of regions in the three-dimensional feature space, along with the generated scoring intervals and constructed personalized rule sets, adapt to the current virtual imaging object, improving rule targeting. Initial scores are calibrated based on the personalized rule set, and targeted adjustments make the optimized capability score more aligned with rule requirements. Multi-dimensional weighted aggregation of the optimized score integrates data from four virtual imaging dimensions, reflecting the different impacts of each dimension on overall capability, forming a comprehensive evaluation result. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating a multimodal intelligent analysis method for assessing the capabilities of power industry personnel, provided by an embodiment of the present invention.
[0014] Figure 2This is a schematic diagram of a multimodal intelligent analysis system for assessing the capabilities of power industry personnel, provided by an embodiment of the present invention. Detailed Implementation
[0015] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0016] like Figure 1 As shown, an embodiment of the present invention proposes a multimodal intelligent analysis method for assessing the capabilities of power industry personnel. The method includes the following steps: Step 1: Based on the preset weight configuration strategy, perform weighted fusion on the multimodal input data that has been time-aligned and formatted to generate a fused multimodal data matrix; extract multimodal feature vectors based on this matrix to characterize operation behavior patterns, interaction response characteristics and state judgment indicators; Step 2: Input the multimodal feature vector into the pre-trained neural network regression model for high-dimensional mapping calculation to obtain the initial capability assessment score; Step 3: Based on the initial ability assessment score, extract typical modal features from speech and action modalities, calculate the difference measure between them and the baseline patterns in the virtual talent template library, and generate feature difference vectors; use the geometric hash algorithm to quickly match and map the feature difference vectors to obtain the appropriate virtual type templates. Step 4: Based on the virtual type template, dynamically divide and merge regions in the three-dimensional feature space composed of operation behavior patterns, state judgment indicators and interactive response characteristics to generate a scoring interval adapted to the current virtual imaging object; construct a personalized scoring adjustment rule set through interval mapping and normalization operations. Step 5: Based on the personalized scoring adjustment rule set, calibrate the initial ability assessment score to obtain the optimized ability score; Step 6: Perform multi-dimensional weighted aggregation on the optimized capability score to generate a comprehensive capability assessment result. The assessment result covers four virtual imaging dimensions: operational behavior pattern, interactive response characteristics, state judgment index, and execution stability.
[0017] In this embodiment of the invention, a preset weight configuration strategy is used to perform weighted fusion of multimodal data after temporal alignment and format normalization. This fully combines the characteristics of different modal data, allowing the fused multimodal data matrix to more comprehensively integrate various input information. Based on this matrix, multimodal feature vectors are extracted, simultaneously covering operational behavior patterns, interaction response characteristics, and state judgment indicators, making the feature representation richer and more aligned with evaluation needs. The multimodal feature vectors are input into a pre-trained neural network regression model for high-dimensional mapping calculation. The pre-trained model can rely on existing training accumulation to reduce dependence on current data. High-dimensional mapping can transform complex multimodal features into initial capability evaluation scores that are more suitable for the evaluation scenario, reducing the complexity of subsequent processing. Based on the initial capability evaluation scores, canonical speech and action modalities are extracted. The system utilizes type features to focus on the core information of key modalities, reducing the interference of redundant data in the evaluation. It calculates the difference between the virtual talent template library's benchmark pattern and generates feature difference vectors, clearly quantifying the differences between features and providing a basis for pattern matching. Geometric hashing algorithms are used to quickly match and map feature difference vectors to patterns, improving matching efficiency. Based on the virtual type template, dynamic region division and merging are performed in the three-dimensional feature space. This three-dimensional feature space can simultaneously associate operational behavior patterns, state judgment indicators, and interactive response characteristics. Dynamic division and merging allow the generated scoring intervals to better reflect the actual situation of the current virtual imaging object. Personalized scoring adjustment rule sets are constructed through interval mapping and normalization operations. Normalization unifies the scale of data from different dimensions, making the construction of the rule set more standardized. The initial capability assessment score is calibrated based on a personalized scoring adjustment rule set. The personalized nature of the rule set can be adjusted to address the adaptation differences between the initial score and the current virtual imaging object, improving the matching degree between the score and the actual situation of the object. The optimized capability score is then aggregated with multi-dimensional weighting, which combines the different importance of four virtual imaging dimensions: operation behavior mode, interaction response characteristics, state judgment indicators, and execution stability. The comprehensive capability assessment result more comprehensively covers the assessment requirements, and the weighting method reflects the assessment weight of each dimension, making the comprehensive result more targeted.
[0018] In a preferred embodiment of the present invention, the method further includes the following step before step 1: The system collects multimodal raw data streams generated during power operation and maintenance, including text records, voice signals, motion sensing sequences, and numerical logs. Preprocessing of these raw data streams yields time-aligned and formatted multimodal data. Specifically, this includes: deploying multi-source acquisition devices and systems to acquire data from various information carriers; extracting text records from operation and maintenance work record systems, task documents, and fault documents, including operation time, location, equipment model, and operation description; acquiring voice signals through portable recording devices or on-site audio pickup devices, including communication instructions, equipment sounds, and anomaly feedback; acquiring motion sensing sequences using inertial measurement units worn by personnel, posture sensors, and motion capture devices on tools, recording limb trajectories, operational force, and action duration; and exporting numerical logs from equipment monitoring systems, sensor networks, and smart meters, including parameters such as voltage, current, temperature, and load rate. First, the original data of each modality is standardized in format, the text is converted into uniformly encoded structured text, the speech is denoised and the sampling rate and format are unified, the motion sensing sequence is adjusted to sample data at the same time interval, and the fields and data types of the numerical log are standardized. Then, based on the job start time, a unified precision timestamp is added to the data of each modality. Time sequence alignment is achieved by timestamp matching to eliminate misalignment caused by differences in device response. Finally, multimodal data with time synchronization and uniform format is obtained.
[0019] Based on the regularized multimodal data, the corresponding work environment level and operation task type are identified. According to the work environment level and operation task type, a fusion weight coefficient is dynamically configured for each modal data, generating a weight allocation scheme. Specifically, this includes: extracting scene-related features from the regularized data; text filtering of work location, equipment voltage level, and environmental keywords; extracting parameter fluctuation range and numerical intervals from numerical logs; speech recognition of high-voltage, outdoor, and other environment-related segments; and action sensor sequence analysis of special environmental protection actions. Referring to preset standards, the environment is divided into two categories: complex (high voltage, severe outdoor weather, multi-equipment collaborative operation) and normal (low voltage, stable indoor environment, single equipment operation). Simultaneously, task-related features are extracted; text is used to obtain task names, objectives, and processes; speech recognition is used for tasks such as inspection and maintenance; and action... The sensor sequence is matched with typical action patterns, such as walking observation during inspection and disassembly and installation during maintenance, to determine the type of operation task. First, a mapping model is established between the environmental level, task type, and the correlation between data of each modality. Combined with historical data, the basic weight range of each modality is determined. Adjustments are made according to the environmental level: in complex environments, the weight of action sensor sequence (associated with operational safety) and numerical log (associated with equipment stability) is increased, while the weight of text and voice is decreased; in normal environments, the weight of text (associated with process compliance) and voice (associated with instruction accuracy) is increased, while the weight of action and numerical log is adjusted. The process is further refined according to the task type: for safety-oriented maintenance tasks, the weight of action and numerical log is further increased; for process-oriented inspection tasks, the weight of text and voice is further increased. The specific weight coefficients of each modality are calculated to generate a scheme that includes the weight ratio of text, voice, action sensor sequence, and numerical log.
[0020] In this embodiment of the invention, multimodal raw data streams are collected during power operation and maintenance, covering text records, voice signals, motion sensor sequences, and numerical logs, which can comprehensively capture information from different dimensions during the operation. The multimodal raw data streams are preprocessed to obtain multimodal data with time alignment and format regularity, which can eliminate the misalignment problem of different modal data in the time dimension and unify the data format standard. Based on the identification of corresponding work environment levels and operation task types using regularized multimodal data, the system can accurately associate actual work scenario attributes with the feature information in the regularized data, giving the weight configuration a definite scenario basis and avoiding blind weight settings that are detached from the actual work situation. The system can dynamically configure the fusion weight coefficients for each modality of data according to the work environment level and operation task type, so that the weight allocation can be adapted to the complexity of the current work environment and the core requirements of the operation task. The generated weight allocation scheme can make the subsequent multimodal data fusion more in line with the information needs of the specific work scenario, and improve the pertinence of the fused data in representing work behavior.
[0021] In a preferred embodiment of the present invention, step 1 above involves weighted fusion of the time-aligned and format-normalized multimodal input data based on a preset weight configuration strategy to generate a fused multimodal data matrix; and extracting multimodal feature vectors based on this matrix to characterize operational behavior patterns, interactive response characteristics, and state judgment indicators, including: Step 11: Based on the weight allocation scheme, perform temporal and spatial alignment processing on the regularized multimodal data to generate a spatiotemporally aligned multimodal data sequence. Specifically, this includes: firstly, extracting the priority and associated scenario information of each modal data in the weight allocation scheme to determine the adjustment basis for each modal data during the alignment process; secondly, using key time nodes of power operation and maintenance as anchor points, including equipment startup time, operation start time, and fault alarm time, uniformly calibrating the operation timestamps in text records, the sampling time axis of voice signals, the acquisition time sequence of action sensor sequences, and the recording time of numerical logs, through interpolation. Alternatively, the time length of each modality data can be adjusted by truncation to ensure that each modality data within the same time segment corresponds to the same work stage. In terms of spatial alignment, the installation location of each data acquisition device is marked by combining the spatial layout map of the power operation site, such as the equipment area corresponding to the microphone, the operating position of the motion sensor wearer, and the physical location of the equipment to which the numerical log belongs. The limb movement trajectory in the motion sensing sequence is associated with the corresponding equipment spatial coordinates, the voice signal is bound with the equipment information of the acquisition location, and the equipment location description in the text record is matched with the actual spatial coordinates, forming a multimodal data sequence that is synchronized in the time dimension and associated in the spatial dimension.
[0022] Step 12: Based on the spatiotemporally aligned multimodal data sequence, perform weighted fusion calculation according to the fusion weight coefficients corresponding to each modality to generate a weighted fused multimodal data matrix. Specifically, this includes: firstly, extracting the fusion weight coefficients corresponding to text records, speech signals, motion sensing sequences, and numerical logs from the weight allocation scheme to determine the contribution ratio of different modal data in the fusion process; quantizing each spatiotemporally aligned modal data separately, converting the job description in the text records into structured feature values, extracting acoustic feature values such as Mel frequency cepstral coefficients from the speech signals, converting the motion sensing sequences into motion feature values such as limb joint angles and motion speeds, and converting parameters such as voltage and current in the numerical logs into standardized operational feature values; according to the weight coefficients of each modality, performing a weighted summation operation on the quantized feature values of each modality within the same time segment, with each time segment corresponding to a set of fused feature values; arranging the fused feature values of all time segments in chronological order, where each row of data represents the fusion feature of a time segment, and each column of data represents a feature dimension, thus constructing a weighted fused multimodal data matrix.
[0023] Step 13 involves performing principal component analysis on the weighted fused multimodal data matrix to obtain the dimensionality-reduced principal component feature set. Specifically, this includes: first, preprocessing the weighted fused multimodal data matrix by calculating the mean and standard deviation of each feature dimension; then, converting the eigenvalues of each dimension into standardized data with a mean of 0 and a variance of 1 using a standardization formula to eliminate the influence of dimensional differences between different feature dimensions; next, calculating the covariance matrix of the standardized data matrix to reflect the degree of linear correlation between different feature dimensions; then, solving for the eigenvalues and corresponding eigenvectors of the covariance matrix and sorting the eigenvectors in descending order of eigenvalues; selecting the top few eigenvectors whose cumulative contribution rate reaches a preset information retention threshold as principal component directions, typically set to retain more than 85% of the original data information; finally, performing matrix multiplication on the standardized multimodal data matrix and the selected principal component directions to project the high-dimensional fused data into the low-dimensional principal component space, obtaining the dimensionality-reduced principal component feature set, with each sample data corresponding to a set of principal component eigenvalues.
[0024] Step 14: Based on the principal component feature set, obtain the final discriminant projection vector through linear discriminant analysis. Specifically, this includes: first, labeling the principal component feature set by category; based on the differences in operational behavior patterns, interactive response characteristics, and status judgment indicators of power operation and maintenance, dividing the samples in the principal component feature set into different categories, such as inspection operation, maintenance operation, and fault diagnosis; calculating the mean vector of each category, which reflects the average level of the corresponding category's samples across each principal component dimension; and simultaneously calculating the overall mean vector of all categories; constructing the intra-class divergence matrix and the inter-class divergence matrix respectively. The scatter matrix is obtained by calculating the sum of squared deviations of samples within each category from the mean vector of that category, reflecting the dispersion of samples within the same category. The inter-class scatter matrix is obtained by calculating the sum of squared deviations of the mean vector of each category from the overall mean vector, reflecting the degree of difference between samples of different categories. The product matrix of the inverse of the intra-class scatter matrix and the inter-class scatter matrix is solved, and the eigenvalue decomposition of the product matrix is performed to obtain the corresponding eigenvalues and eigenvectors. The eigenvectors are selected in descending order of eigenvalues, and the eigenvectors with the largest eigenvalues are selected to form a projection matrix, which is the final discriminant projection vector.
[0025] Step 15: Perform spatial projection transformation on the principal component feature set using the final discriminant projection vector to extract the feature combination with the maximum class discrimination, forming a multimodal feature vector for characterizing operational behavior patterns, interaction response characteristics, and state judgment indicators. Specifically, this includes: first, performing matrix operations on each sample feature vector in the principal component feature set and the final discriminant projection vector to complete the projection transformation of the sample data from the principal component space to the discriminant space, obtaining the projected feature vector of each sample in the discriminant space; then, performing statistical analysis on all projected sample feature vectors, calculating the mean difference and variance difference of different categories of samples in each projected feature dimension, and selecting projected feature vectors with large inter-category mean differences and small intra-category variance differences. Feature dimensions are selected, and the features corresponding to these dimensions have the highest class discriminative power. Based on the selected feature dimensions, feature values of each sample are extracted in these dimensions. The feature values are grouped according to the classification logic of operation behavior mode, interaction response characteristics, and state judgment indicators. The feature dimensions corresponding to operation behavior mode include motion trajectory related projection features, the feature dimensions corresponding to interaction response characteristics include voice command response related projection features, and the feature dimensions corresponding to state judgment indicators include numerical log parameter change related projection features. The grouped feature values are combined in a fixed order to form a multimodal feature vector for each sample. Each dimension of this vector corresponds to the representation information of operation behavior mode, interaction response characteristics, and state judgment indicators, respectively.
[0026] In this embodiment of the invention, combining a weight allocation scheme for temporal and spatial alignment allows the alignment operation to align with the current working environment and task requirements. Temporal alignment reinforces the time synchronization of data across modalities, while spatial alignment supplements the spatial location of data and its relevance to the operational scenario. The combination of both improves the spatiotemporal matching degree of multimodal data sequences and reduces fusion information deviation caused by spatiotemporal misalignment. Weighted fusion based on the fusion weight coefficients of each modality ensures that key modalities occupy a reasonable proportion in the data matrix, avoiding the dilution of critical information through equal fusion. Weight-guided fusion calculation makes the data matrix reflect the importance of each modality in the current scenario, and the matrix information better matches the representation requirements of operational behavior patterns, interactive response characteristics, and state judgment indicators. Principal component analysis is used... Dimensionality reduction analysis of high-dimensional fused data matrices can eliminate redundant and repetitive features. During dimensionality reduction, key feature information is preserved, allowing the generated principal component feature set to simplify the data structure while still carrying the core content of multimodal data. Using linear discriminant analysis to process the principal component feature set allows focusing on the differences in information across different categories within the feature set, finding the projection direction that maximizes category discrimination. Spatial projection transformation strengthens features in the principal component feature set that conform to the discriminant projection vector direction, while weakening irrelevant features with low category discrimination. The feature combination with the highest category discrimination is extracted to form a vector, which accurately captures differences in different operational behavior patterns, interaction response characteristics, and state judgment indicators, making the representation of these three aspects more targeted and effective.
[0027] In a preferred embodiment of the present invention, step 2 above, which involves inputting the multimodal feature vector into a pre-trained neural network regression model for high-dimensional mapping calculation to obtain an initial capability assessment score, includes: Step 21: Based on the multimodal feature vector, the Z-score standardization method is used to perform feature scaling to obtain a standardized feature vector. Specifically, this includes: first, calculating the mean and standard deviation of each feature dimension in the multimodal feature vector, where the mean is obtained by summing all feature values in that dimension and dividing by the number of feature values, and the standard deviation is obtained by averaging the squares of the differences between each feature value and the mean in that dimension and then taking the square root; then, for each feature value in the multimodal feature vector, subtracting the mean of its dimension from the feature value and dividing by the standard deviation of that dimension to complete the standardization of a single feature value; performing the above operation sequentially on the feature values of all dimensions in the multimodal feature vector to finally obtain a standardized feature vector with uniform scale for each feature dimension.
[0028] Step 22: Input the standardized feature vector into the pre-trained neural network regression model; perform forward propagation calculation through the input layer, hidden layer, and output layer of the neural network regression model to obtain a multi-dimensional initial capability evaluation vector from the output layer. Specifically, the process of inputting the standardized feature vector into the pre-trained neural network regression model and obtaining the multi-dimensional initial capability evaluation vector through forward propagation calculation of each layer is as follows: First, construct and train the pre-trained neural network regression model, and then perform forward propagation calculation: Determine the overall structural parameters of the model, which includes an input layer, a hidden layer, and an output layer; the number of neurons in the input layer is consistent with the dimension of the standardized multimodal feature vector to ensure that all information in the feature vector can be fully received; the hidden layer is set with a multi-layer fully connected structure, and the number of neurons in each layer is adapted and adjusted according to the complexity of the power operation and maintenance operation evaluation scenario, and each layer is equipped with a ReLU activation function to enhance the model's ability to fit nonlinear features; the number of neurons in the output layer is set to 3, corresponding to the three evaluation dimensions of operation behavior mode, interaction response characteristics, and state judgment index, to achieve synchronous representation of capabilities in different dimensions.
[0029] Multimodal feature vectors with corresponding capability assessment results labeled in historical power operation and maintenance data were selected as the training dataset, while a portion of the data was used as the validation set. The mean squared error loss function was used to calculate the error between the model's predictions and the labeled results to measure the accuracy of the model's output. The Adam optimizer was used to iteratively update the model parameters. In each iteration, the feature vectors from the training set were first input into the model for forward propagation to obtain the predicted values. Then, the gradient of the loss function with respect to the weights and biases of each layer was calculated using the backpropagation algorithm. The weights and biases were adjusted based on the gradient information to reduce the prediction error. The validation set was continuously used to monitor model performance during the iteration process. An early stopping strategy was adopted to avoid overfitting. Training was stopped when the error on the validation set no longer decreased after several consecutive iterations. Finally, a pre-trained neural network regression model adapted to the power operation and maintenance capability assessment scenario was obtained.
[0030] The standardized feature vectors are input into the input layer of the pre-trained neural network regression model. Each neuron in the input layer receives the corresponding feature value from the feature vector and passes the complete feature value to the first hidden layer. The first hidden layer performs linear calculations on the received feature data using a preset weight matrix and bias term. Then, it performs a non-linear transformation on the linear calculation result using the ReLU activation function to extract complex features from the data. The transformed data is then passed to the next hidden layer. Subsequent hidden layers repeat the linear calculation and non-linear transformation operations, gradually deepening the processing of feature information, until the last hidden layer passes the processed feature data to the output layer. The output layer integrates the data passed from the last hidden layer through linear calculations, generating values corresponding to three dimensions: operational behavior pattern, interaction response characteristics, and state judgment index. These three values are combined in order of evaluation dimensions to form a multi-dimensional initial capability evaluation vector.
[0031] Step 23: Based on the initial capability assessment vector, a nonlinear transformation is performed using the Sigmoid activation function to obtain the transformed vector data. Specifically, this includes: first, obtaining the value of each dimension in the initial capability assessment vector. The Sigmoid activation function can map different ranges of input values to the interval between 0 and 1 through a specific nonlinear calculation method. At the same time, it can strengthen the representation of the complex nonlinear relationship between the assessment dimensions (operation behavior mode, interaction response characteristics, and state judgment indicators) in the initial capability assessment vector, making the vector data better reflect the correlation characteristics between the dimensions. Then, each dimension value of the initial capability assessment vector is substituted into the Sigmoid activation function for calculation, and the Sigmoid transformation of all dimensions of the initial capability assessment vector is completed in sequence. The resulting transformed vector data, where the values of each dimension are all in the interval between 0 and 1, serves two purposes: firstly, it meets the data range requirements of subsequent linear scaling calculations; secondly, it provides a suitable data foundation for obtaining a standardized output vector based on linear scaling and then calculating the initial capability assessment score. At the same time, it makes the vector data more likely to reflect the relative differences of each assessment dimension in subsequent processing.
[0032] Step 24: Based on the transformed vector data, map it to a preset standard numerical range through linear scaling calculation to obtain a standardized output vector. Specifically, this includes: first, pre-setting a standard numerical range (e.g., 0 to 100 points) that meets the requirements of power operation and maintenance personnel competency assessment; calculating the difference between the upper and lower limits of this range to obtain the range span; then determining the current numerical range of the transformed vector data; calculating the difference between the maximum and minimum values of the current range to obtain the current span; for each dimension value in the transformed vector data, subtracting the minimum value of the current range from the value, multiplying by the ratio of the standard range span to the current span, and finally adding the lower limit of the standard range to complete the linear scaling of a single dimension value; performing the above scaling operation on all dimension values of the vector to obtain a standardized output vector where each dimension value is within the preset standard range.
[0033] Step 25: Based on the standardized output vector and preset weight coefficients corresponding to the operation behavior mode, interaction response characteristics, and status judgment indicators, a weighted summation algorithm is used to calculate the weighted sum, which is the initial capability assessment score. Specifically, this includes: first, based on the importance requirements of power operation and maintenance work for each assessment dimension, preset weight coefficients corresponding to the operation behavior mode, interaction response characteristics, and status judgment indicators, and the sum of the weight coefficients of the three dimensions is 1; extract the values corresponding to the three assessment dimensions from the standardized output vector, multiply the value of each dimension by the preset weight coefficient of that dimension to obtain the weighted score of each dimension; add the weighted scores of the three dimensions together, and the resulting weighted sum is the initial capability assessment score that can comprehensively reflect the situation of the three assessment dimensions.
[0034] In this embodiment of the invention, Z-score standardization is used for feature scaling, which unifies the scale of multimodal feature vectors; avoids interference with model calculation due to differences in the numerical range of features of different dimensions, and provides a stable and consistent feature foundation for subsequent model input; inputting the standardized feature vectors into the pre-trained neural network allows the pre-trained model to efficiently process features based on existing accumulation; through forward propagation of the input layer, hidden layer, and output layer, the correlation between features can be fully explored; the generated multidimensional initial capability assessment vector can carry assessment-related information from multiple aspects; using the Sigmoid activation function for nonlinear transformation can map the vector data to a reasonable numerical range; enhance the nonlinear expressive ability of the data and adapt to the nonlinear correlation requirements between features and scores in capability assessment; mapping the transformed vector to a preset standard interval through linear scaling allows the output vector values to conform to the numerical specifications of the assessment scenario; facilitates subsequent calculation and result comparison, and maintains the consistency of the initial assessment data; combined with the preset weights of the corresponding indicators, the initial score is calculated by weighted summation, which can reflect the different importance of operation behavior patterns, interaction response characteristics, and state judgment indicators in the assessment; and makes the sum result more in line with the assessment focus, making the initial capability assessment score more targeted.
[0035] In a preferred embodiment of the present invention, step 3 above involves extracting typical modal features from speech and action modalities based on the initial ability assessment score, calculating the difference measure between these features and the baseline patterns in the virtual talent template library, and generating a feature difference vector. A geometric hash algorithm is then used to quickly match and map the feature difference vector to obtain a suitable virtual type template, including: Step 31: Based on the initial capability assessment score, a subset of key features is selected from the raw feature data of the speech and action modalities. Specifically, this includes: first, obtaining the initial capability assessment score and determining the key points of power operation and maintenance work assessment corresponding to the score; for example, if the score is low, the focus is on features related to operational standardization, and if the score is high, the focus is on features related to operational efficiency. Then, the raw feature data of the speech modalities is extracted, including the speech rate of the operator's communication instructions, the accuracy of the instructions, the frequency and intensity of equipment operation sounds, and the recognition rate of abnormal sounds. At the same time, the raw feature data of the action modalities is extracted, including the amplitude of the operator's limb operations, the continuity of the movements, the conformity of standard movements, and the force and angle of tool use. Based on the assessment focus determined by the initial capability assessment score, features that have a significant impact on the assessment results are selected; for example, if the score is low, the features of movement continuity and instruction accuracy are selected, and if the score is high, the features of operation amplitude and tool use angle are selected, thus forming a subset of key features of the speech and action modalities.
[0036] Step 32: Based on the key feature subset, a distance metric algorithm is used to calculate the degree of difference between it and each benchmark pattern in the virtual talent template library to obtain the initial difference measurement result. Specifically, this includes: first, determining the types of benchmark patterns stored in the virtual talent template library, including standard voice patterns and standard action patterns corresponding to different power operation and maintenance tasks (such as inspection, maintenance, troubleshooting, and equipment debugging). Each benchmark pattern contains feature items consistent with the dimensions of the key feature subset; then, selecting a suitable distance metric algorithm, such as the Euclidean distance algorithm or the Manhattan distance algorithm; matching each feature item in the key feature subset with the feature items of the corresponding dimensions of each benchmark pattern in the virtual talent template library, calculating the difference value of the two in each dimension using the distance metric algorithm, and integrating the difference values of all dimensions to obtain the initial difference measurement result between each benchmark pattern and the key feature subset, such as the initial difference measurement result between the inspection benchmark pattern and the current key feature subset, the initial difference measurement result between the maintenance benchmark pattern and the current key feature subset, etc.
[0037] Step 33: Based on the initial difference measurement results, dimensionality reduction is performed using principal component analysis to obtain the dimensionality-reduced difference features. Specifically, this includes: first, counting the number of dimensions in the initial difference measurement results obtained in step 32, which is consistent with the number of benchmark patterns in the virtual talent template library. If the template library contains 20 benchmark patterns, then the initial difference measurement results contain 20 dimensions; then, starting the principal component analysis method, using the data of each dimension of the initial difference measurement results as input, and by calculating the covariance matrix, solving for eigenvalues and eigenvectors, selecting principal components whose explained variance exceeds a preset threshold (e.g., 85%); projecting the initial difference measurement results onto the eigenvector space corresponding to the selected principal components to obtain the dimensionality-reduced difference features with fewer dimensions, such as reducing the initial 20-dimensional results to 5-dimensional difference features, thus eliminating redundant information between the difference data of different benchmark patterns.
[0038] Step 34: Based on the dimensionality-reduced difference features, cluster analysis is used to classify patterns and identify the closest baseline pattern category. This includes: first, determining the type of cluster analysis method, such as the K-means clustering algorithm; and determining the number of clusters based on the actual classification of baseline patterns in the virtual talent template library (e.g., inspection, maintenance, and troubleshooting). The dimensionality-reduced difference features obtained in Step 33 are used as cluster input data. Cluster centers are initialized, and the distance between each difference feature data point and each cluster center is calculated. Each difference feature data point is assigned to the nearest cluster based on the distance, and the cluster centers are updated. The distance calculation and assignment process is repeated until the cluster centers no longer change. After clustering, the baseline pattern category corresponding to each cluster is analyzed. The correlation between the current dimensionality-reduced difference feature's cluster and each baseline pattern category is calculated, and the closest baseline pattern category (i.e., the one with the highest correlation) is identified. For example, the current difference feature is most closely related to the maintenance baseline pattern category.
[0039] Step 35: Based on the baseline pattern category, extract the corresponding baseline feature vector and calculate the difference between the key feature subset and the baseline feature vector in each dimension. Specifically, this includes: first, retrieving all baseline patterns identified in step 34 that are closest to the baseline pattern category from the virtual talent template library; extracting the feature vectors of all baseline patterns under this category and calculating the average vector; using this average vector as the baseline feature vector of this baseline pattern category; the dimension of this vector is consistent with the dimension of the key feature subset, such as if the key feature subset contains 5 dimensions, then the baseline feature vector also contains 5 dimensions; then comparing the corresponding dimensions of the key feature subset and the baseline feature vector one by one, such as comparing the value of the action coherence dimension in the key feature subset with the value of the action coherence dimension in the baseline feature vector, and comparing the value of the instruction accuracy dimension in the key feature subset with the value of the instruction accuracy dimension in the baseline feature vector; calculating the difference between the value of the key feature subset and the value of the baseline feature vector in each corresponding dimension to obtain the difference data for each dimension, such as the difference in the action coherence dimension being 0.12 and the difference in the instruction accuracy dimension being 0.08.
[0040] Step 36: Based on the differences in each dimension, a standardized feature difference vector is generated through normalization processing. Specifically, this includes: first, determining the target interval for normalization processing, which is set to 0 to 1 or -1 to 1 according to the numerical range requirements of power operation and maintenance capability assessment; statistically analyzing the maximum and minimum values of the differences in each dimension obtained in step 35, and calculating the value range of the difference in each dimension; mapping the difference data of each dimension to the preset target interval through a linear transformation formula, such as if the original range of the difference in a certain dimension is -0.5 to 0.3 and the target interval is 0 to 1, then each difference in that dimension is proportionally converted to a value between 0 and 1; integrating the difference data of all dimensions after normalization processing to generate a standardized feature difference vector in which the values of each dimension are within the target interval, ensuring that the difference data of different dimensions have a unified dimension.
[0041] Step 37: Based on the feature difference vector, perform dimensionality reduction using principal component analysis to obtain low-dimensional features. Specifically, this includes: first, analyzing the dimensionality of the standardized feature difference vector generated in step 36; if the vector contains 3 dimensions for speech modality and 5 dimensions for action modality, totaling 8 dimensions; then applying principal component analysis, inputting the 8-dimensional data of the standardized feature difference vector into the analysis model, calculating the covariance between the data of each dimension, and determining the proportion of variance explained by each principal component; selecting principal components that can cover the main information according to a preset information retention threshold (e.g., 90%), such as selecting the first 3 principal components, whose total explained variance reaches more than 90%; projecting the standardized feature difference vector onto the space corresponding to the selected principal components to obtain low-dimensional features with reduced dimensionality, such as reducing the 8-dimensional vector to 3-dimensional features, thus reducing the computational complexity of subsequent data processing.
[0042] Step 38: Based on the low-dimensional features, cluster analysis is used to classify patterns and identify the similarity with each benchmark pattern in the virtual talent template library. Specifically, this includes: first, preprocessing all benchmark patterns in the virtual talent template library by standardizing and reducing the dimensionality of the feature vector of each benchmark pattern according to the methods in steps 36 and 37 to obtain the low-dimensional features corresponding to each benchmark pattern and establishing the correspondence between benchmark patterns and low-dimensional features; then, selecting hierarchical clustering algorithm as the cluster analysis method, and inputting the low-dimensional features of the data to be evaluated obtained in step 37 and the low-dimensional features of all benchmark patterns into the clustering model; calculating the Euclidean distance or cosine distance between the low-dimensional features of the data to be evaluated and the low-dimensional features of each benchmark pattern, with a smaller distance indicating a higher similarity; converting the calculated distance value into the corresponding similarity value, such as a similarity of 0.8 when the distance is 0.2 and a similarity of 0.5 when the distance is 0.5, thereby identifying the similarity between the data to be evaluated and each benchmark pattern in the virtual talent template library.
[0043] Step 39: Based on the similarity, select the top K benchmark patterns with the highest similarity as a candidate template set. Specifically, this includes: first, determining the K value according to the evaluation accuracy requirements of power operation and maintenance work and the size of the virtual talent template library. For example, the K value is set to 10 when the template library contains 50 benchmark patterns, and the K value is set to 5 when the template library contains 30 benchmark patterns; sorting the similarity values of the data to be evaluated and each benchmark pattern obtained in Step 38 in descending order to form a similarity ranking list; starting from the first position of the ranking list, selecting the benchmark patterns corresponding to the top K similarity values in sequence, such as selecting the top 10 benchmark patterns; integrating these K benchmark patterns to form a candidate template set, which contains K benchmark patterns with the highest similarity to the data to be evaluated, thus narrowing the scope of subsequent template matching.
[0044] Step 310: Based on the candidate template set, extract the corresponding baseline feature vectors and construct a hash table structure. Specifically, this includes: first, retrieving the complete baseline feature vector of each candidate template in the candidate template set determined in step 39 from the virtual talent template library. This vector contains all key feature dimensions of the speech modality and action modality, such as speech rate, instruction accuracy, and voice recognition rate, and the amplitude, continuity, force, and angle of the action. Then, designing the structure of the hash table, using the unique identifier of the candidate template as the key of the hash table, such as inspection template 001, maintenance template 003, etc., and using the complete baseline feature vector corresponding to each candidate template as the value of the hash table. According to the designed structure, storing the identifier of each candidate template and its corresponding baseline feature vector into the hash table one by one, completing the construction of the hash table, so that subsequent queries of candidate template features can be quickly located by identifier.
[0045] Step 311: Based on the hash table structure, a geometric hash algorithm is used for fast matching to calculate the matching degree between the feature difference vector and each candidate template. Specifically, this includes: first, reading the hash table constructed in step 310 to obtain the identifiers of all candidate templates and their corresponding baseline feature vectors stored therein; then, starting the geometric hash algorithm to convert the standardized feature difference vector generated in step 36 into a set of feature points in geometric space, and simultaneously converting the baseline feature vector of each candidate template in the hash table into a corresponding set of feature points; using the hash function of the geometric hash algorithm, quickly matching the feature point set of the data to be evaluated with the feature point set of the candidate templates in the hash table, and calculating parameters such as the overlap and distance deviation between the two in geometric space; converting these parameters into corresponding matching degree values, such as a matching degree of 0.8 when the overlap is 80% and a matching degree of 0.9 when the distance deviation is 0.1, thereby obtaining the matching degree between the feature difference vector and each candidate template.
[0046] Step 312: Based on the matching degree, select the candidate template with the highest matching degree as the adapted virtual type template. Specifically, this includes: first, collecting the matching degree values of the feature difference vector calculated in step 311 and each candidate template to form a matching degree list; comparing all values in the matching degree list to find the candidate template corresponding to the largest matching degree value; confirming the virtual talent type corresponding to the candidate template, such as the template belonging to the senior power inspection talent template, intermediate power maintenance talent template, etc.; and determining the candidate template as the virtual type template adapted to the current data to be evaluated.
[0047] In this embodiment of the invention, selecting key feature subsets of speech and action modalities based on initial capability assessment scores reduces redundant feature interference and focuses on core features relevant to the assessment. Using a distance metric algorithm to calculate the degree of difference between the key feature subsets and each benchmark pattern quantifies the difference, providing data for judging feature matching and avoiding subjective judgment bias. Principal component analysis is performed on the initial difference measurement results to reduce dimensionality, eliminating redundant information in the difference data, simplifying the data structure, and retaining key difference features, thus reducing the complexity of subsequent processing. Cluster analysis is used to classify the dimensionality-reduced difference features, quickly identifying the closest benchmark pattern category, narrowing the matching range, and improving the efficiency of subsequent pattern matching. Extracting the benchmark feature vector corresponding to the benchmark pattern category and calculating the difference between the key feature subset and each dimension refines the difference representation. Normalizing the differences in each dimension to generate feature difference vectors unifies the difference scale and eliminates data magnitude differences across dimensions. The impact of differences: Principal component analysis (PCA) is used to reduce the dimensionality of the feature difference vectors to obtain low-dimensional features, which further simplifies the data dimensionality, reduces the computational load of subsequent pattern classification, and improves processing speed. Cluster analysis is used to classify low-dimensional features and identify their similarity to each benchmark pattern, which can quickly quantify the matching degree between features and benchmark patterns, providing a basis for candidate template selection. Selecting the top K benchmark patterns with the highest similarity as the candidate template set can focus on high-matching templates and narrow the final adaptation range. Extracting the benchmark feature vectors of the candidate template set to construct a hash table can optimize the data storage structure and provide an efficient data query foundation for subsequent fast matching. Using the geometric hash algorithm based on the hash table for fast matching, the matching degree between the feature difference vectors and each candidate template is calculated, improving the matching speed and shortening the template adaptation time. Selecting the candidate template with the highest matching degree as the adaptation virtual type template can ensure the adaptability of the template to the current features and provide an accurate basis for subsequent scoring interval division.
[0048] In a preferred embodiment of the present invention, step 4 above involves dynamically dividing and merging regions in a three-dimensional feature space composed of operation behavior patterns, state judgment indicators, and interaction response characteristics, based on the virtual type template, to generate a scoring interval adapted to the current virtual imaging object; and constructing a personalized scoring adjustment rule set through interval mapping and normalization operations, including: Step 41: Based on the virtual type template, construct a three-dimensional feature space consisting of three dimensions: operation behavior mode, state judgment index, and interaction response characteristics. Specifically, this includes: first, extracting all feature parameters related to the operation behavior mode, state judgment index, and interaction response characteristics from the virtual type template. The feature parameters for the operation behavior mode include indicators such as the standardization of the action sequence of the virtual imaging object in the power operation and maintenance simulation and the completeness of the operation process; the feature parameters for the state judgment index include indicators such as the accuracy of identifying the equipment operating status and the response speed to abnormal situations; and the feature parameters for the interaction response characteristics include indicators such as the degree of coordination with other elements in the simulated operation scenario and the timeliness of instruction execution. Then, determine the value range for each dimension based on the feature parameters in the virtual type template. Based on the standard threshold and the fluctuation range of characteristic data of similar historical virtual imaging objects, the value range of the operation behavior mode dimension is set to 0 to 100 (corresponding to the range from extremely non-standard operation to completely standard operation), the value range of the state judgment index dimension is set to 0 to 100 (corresponding to the range from completely wrong state judgment to completely accurate judgment), and the value range of the interaction response characteristic dimension is set to 0 to 100 (corresponding to the range from extremely untimely response to completely timely response). Finally, with the operation behavior mode as the X-axis, the state judgment index as the Y-axis, and the interaction response characteristic as the Z-axis, the scale interval (each scale corresponds to 1 unit of feature value) and unit label of each coordinate axis are set to ensure that all feature data of the three dimensions in the virtual type template can find the corresponding spatial position on the coordinate axis, thereby constructing a three-dimensional feature space.
[0049] Step 42: Based on the three-dimensional feature space, cluster analysis is used to divide the initial region, resulting in multiple initial feature regions. Specifically, this includes: First, collecting multimodal feature data generated by the current virtual imaging object during power operation and maintenance simulation within the three-dimensional feature space. This data covers action sensor sequence analysis results in the operational behavior mode dimension, numerical log analysis data in the state judgment index dimension, and the association information between voice signals and text records in the interaction response characteristic dimension. All data has been time-aligned and formatted. Then, the key parameters for cluster analysis are determined. Referring to the clustering results of the three-dimensional feature data of similar virtual imaging objects in the past, the initial range of the number of clusters is set to 5 to 8 (depending on the complexity of the current virtual imaging object's operation). (Adjusting the degree appropriately), Euclidean distance is selected as the index to measure the similarity between feature data, and a clustering convergence threshold is set (the iteration stops when the movement distance of the cluster center in two iterations is less than the threshold). Then, the sorted multimodal feature data is input into the clustering analysis tool. The tool calculates the Euclidean distance between any two data points according to the coordinate position of the feature data in three-dimensional space, and groups the data points that are close to each other (less than the preset similarity threshold) into the same group. After the clustering analysis tool completes the iterative calculation, each data group corresponds to an initial feature region. At the same time, the X-axis, Y-axis, and Z-axis boundary range of each initial feature region, as well as the number of feature data points contained in the region and the coordinate information of each data point, are recorded to obtain multiple initial feature regions.
[0050] Step 43: Based on the initial feature regions, and using the feature distribution pattern defined in the virtual type template as the merging basis, perform a merging operation on regions with similar and adjacent feature distributions in space to obtain an optimized feature region division. Specifically, this includes: first, deeply analyzing the virtual type template to extract the feature distribution patterns of operational behavior patterns, state judgment indicators, and interactive response characteristics; these patterns constitute the core content of the feature distribution pattern; then, extracting the feature distribution parameters of each initial feature region, including the mean and variance of the feature values in each of the three dimensions, and the correlation coefficient between the three-dimensional feature values; finally, determining whether adjacent initial feature regions meet the merging conditions. First, check if the two regions share a boundary in three-dimensional space. Then, calculate the difference in feature distribution parameters. If the difference in the mean and variance of the three dimensions is less than the corresponding preset threshold, and the difference in the correlation coefficient is less than the preset threshold, then the feature distribution is determined to be similar. Perform a merging operation on the initial regions that are adjacent and have similar feature distributions. After merging, recalculate the region boundary range (take the minimum value of the coordinate axis boundary of the two regions as the lower limit and the maximum value as the upper limit) and feature distribution parameters (take the weighted mean and weighted variance of the feature data of the two regions). Repeat the above adjacent region judgment and merging operation until there are no regions that meet the merging conditions, and finally obtain the optimized feature region division.
[0051] Step 44: Based on the optimized feature region division, extract the feature distribution parameters of each region to obtain a dynamic scoring interval adapted to the current virtual imaging object. Specifically, this includes: firstly, for each optimized feature region, extracting feature distribution parameters in three dimensions: operation behavior mode, state judgment index, and interaction response characteristics; extracting the minimum, maximum, median, and 95th percentile of feature values in the operation behavior mode dimension; extracting the qualified threshold, excellent threshold, and peak value corresponding to the feature value distribution density within the region according to the virtual type template in the state judgment index dimension; and extracting the feature value fluctuation range, the feature value corresponding to the average response time, and the stable response feature value interval in the interaction response characteristics dimension.
[0052] Then, scoring intervals are constructed based on the feature distribution parameters of each dimension. For the operational behavior mode dimension, the minimum value to the median is set as the basic scoring interval, the median to the 95th percentile as the good scoring interval, and the 95th percentile to the maximum value as the excellent scoring interval. For the state judgment index dimension, the interval below the qualified threshold is set as the interval to be improved, the qualified threshold to the excellent threshold is set as the qualified scoring interval, and the interval above the excellent threshold is set as the excellent scoring interval. For the interaction response characteristics dimension, the interval with a fluctuation range greater than the preset threshold is set as the unstable scoring interval, the fluctuation range meets the requirements and the feature value is within the stable response interval is set as the stable scoring interval, and the feature value corresponding to the average response time within the stable response interval is set as the efficient scoring interval. Finally, the scoring intervals of the three dimensions are integrated to match the regional feature distribution parameters to form a dynamic scoring interval adapted to the current virtual imaging object, and the three-dimensional feature region identifier corresponding to each dynamic scoring interval is recorded.
[0053] Step 45: Based on the dynamic scoring interval, a linear mapping algorithm is used to map the original scoring data to the corresponding interval range to obtain the interval mapping result. Specifically, this includes: first, collecting the original scoring data of the current virtual imaging object's power operation and maintenance work evaluation, including the action specification score of the operation behavior mode, the equipment status recognition score of the status judgment index, and the collaborative response score of the interactive response characteristics. Each dimension's original score has a corresponding value range (e.g., 0 to 100 points); then, for each dimension's original score, by querying the correlation between the dynamic scoring interval and the three-dimensional feature region identifier, finding the three-dimensional feature data to which it belongs. First, define the upper and lower limits of the dynamic scoring interval for each dimension within a given feature region. Then, calculate the linear mapping parameters, setting the original scoring range as [original lower limit, original upper limit] and the dynamic scoring interval range as [dynamic lower limit, dynamic upper limit]. The mapping ratio is (dynamic upper limit - dynamic lower limit) / (original upper limit - original lower limit), and the mapping offset is dynamic lower limit - original lower limit × mapping ratio. Next, substitute the original scores for each dimension into the mapping relationship (mapped score = original score × ratio + offset) to ensure the result falls within the corresponding dynamic interval. Finally, summarize the mapping results for the three dimensions to obtain the interval mapping result.
[0054] Step 46: Based on the interval mapping results, perform data standardization processing using a normalization method to form a personalized scoring adjustment rule set. Specifically, this includes: first, determining the normalization target range as 0 to 100 points (referencing commonly used scales in power industry talent assessment and covering all possible values for the dimension mapping results); then, for each dimension of the interval mapping results, extracting the maximum and minimum values of the mapping results for operational behavior patterns, state judgment indicators, and interactive response characteristics; next, performing normalization calculations on the mapping results of each dimension (normalized value = (mapping result value - minimum value of that dimension) / (minimum value of that dimension)). (Maximum value - minimum value) × 100) to ensure the result falls within the range of 0 to 100; then analyze the difference between the normalized data and the standard data of the corresponding dimension of the virtual type template. For example, if the normalized data of the operation behavior mode is 5 points lower than the template standard, it is determined that the weight of this dimension needs to be increased in subsequent calibration. If the state judgment index is 3 points higher than the standard, it is determined that the weight can be appropriately reduced. Based on this, the scoring adjustment rules for each dimension (including adjustment direction, magnitude and triggering conditions) are formulated; finally, all dimension adjustment rules are integrated to determine the corresponding dimension, adjustment method and triggering conditions of each rule, forming a personalized scoring adjustment rule set adapted to the current virtual imaging object.
[0055] In this embodiment of the invention, a three-dimensional feature space containing operational behavior patterns, state judgment indicators, and interactive response characteristics is constructed. This integrates the three core evaluation dimensions into a single analytical framework, making feature distribution more intuitive. It provides a clear spatial carrier for subsequent region division and merging, facilitating accurate association with virtual type template features. Using cluster analysis for initial region division automatically generates multiple initial regions based on feature similarity, ensuring the initial regions better match the actual feature distribution within the three-dimensional feature space. Merging similar adjacent regions according to the feature distribution pattern of the virtual type template allows the merged regions to better match the feature rules defined by the template. This reduces redundant and mismatched regions, making feature region division more accurate and adaptable. The system first identifies the virtual imaging object; then extracts regional feature distribution parameters to obtain dynamic scoring intervals, allowing these intervals to be generated based on the actual feature distribution of the current object. This avoids the lack of universality of fixed scoring intervals, making the intervals more closely match the feature attributes of the current virtual imaging object. A linear mapping is used to map the original score to the corresponding interval, ensuring that the original score accurately matches the range of the dynamic scoring interval. This ensures the rationality of the scoring data within the interval, avoiding subsequent processing deviations caused by mismatches between the original score and the interval. A personalized scoring adjustment rule set is formed through normalization, unifying the data scale and ensuring consistent adjustment standards within the rule set. Simultaneously, the personalized rule set can be adapted to the current virtual imaging object, providing a targeted basis for subsequent scoring calibration.
[0056] In a preferred embodiment of the present invention, step 5 above, which calibrates the initial ability assessment score based on the personalized scoring adjustment rule set to obtain an optimized ability score, includes: Step 51: Based on the personalized scoring adjustment rule set and the initial capability assessment score, a rule matching algorithm is used to match the scoring data with the adjustment rules to obtain the rule matching results for each dimension. Specifically, this includes: first, obtaining the personalized scoring adjustment rule set, which formulates rules for three dimensions: operational behavior mode, state judgment index, and interaction response characteristics, clarifying the adaptation conditions corresponding to different initial scoring ranges. For example, in the operational behavior mode dimension, 45 to 65 points correspond to basic calibration rules, and 65 to 85 points correspond to advanced calibration rules; in the state judgment index dimension, a score below 55 points corresponds to compensation calibration rules, and a score above 75 points corresponds to fine-tuning calibration rules. The fluctuation of the interaction response characteristic dimension within 3 points corresponds to the stable calibration rule, and the fluctuation between 3 and 8 points corresponds to the dynamic calibration rule. At the same time, the specific values of the three dimensions are extracted from the initial capability assessment score, such as the operation behavior mode 58 points, the state judgment index 52 points, and the interaction response characteristic 73 points. Then, the rule matching algorithm is started to match the dimensions in sequence: the operation behavior mode 58 points conforms to the basic calibration rule of 45 to 65 points, the state judgment index 52 points conforms to the compensation calibration rule below 55 points, and the interaction response characteristic 73 points conforms to the stable calibration rule with fluctuation within 3 points after calculating the fluctuation range. The rule information of each dimension is recorded, and finally, the rule matching results of each dimension are integrated to form the rule matching results of each dimension.
[0057] Step 52: Based on the rule matching results, a difference measurement algorithm is used to calculate the deviation between the initial capability assessment score and the corresponding threshold of the personalized scoring adjustment rule set, obtaining the scoring deviation data. Specifically, this includes: first, extracting the parameters of the matching rules for each dimension from the rule matching results, such as the standard score lower limit of 48 points, upper limit of 68 points, and benchmark score of 55 points for the basic calibration rule of operational behavior mode; the qualified threshold of 55 points and the allowable deviation range of ±4 points for the compensation calibration rule of state judgment index; and the benchmark score of 72 points and the allowable fluctuation range of ±3 points for the stable calibration rule of interactive response characteristics. Then, the values of each dimension of the initial capability assessment score (operational behavior mode 58 points, state judgment index 52 points, interactive response characteristics 73 points) are compared with the corresponding rule parameters and input into the difference measurement algorithm.
[0058] In the operational behavior pattern dimension, the absolute deviation of 58 points from the benchmark score of 55 points is calculated to determine whether 58 points falls within the range of 48 to 68 points, and the deviation value and the status within the range are recorded. In the status judgment index dimension, the absolute deviation of 52 points from the passing threshold of 55 points is calculated to determine whether it is within the allowable range of ±4 points, and the degree of deviation and whether it exceeds the range are recorded. In the interaction response characteristic dimension, the fluctuation deviation of 73 points from the benchmark score of 72 points is calculated to determine whether it is within the range of ±3 points, and the fluctuation deviation is recorded. Finally, the absolute deviation values, whether it exceeds the allowable range, and the status within the range of the three dimensions are integrated to form scoring deviation data containing details of deviations in each dimension.
[0059] Step 53: Based on the scoring deviation data, the initial capability assessment score is calibrated using an interpolation compensation algorithm to obtain the calibrated scoring result. Specifically, this includes: first, extracting deviation information for each dimension from the scoring deviation data in Step 52, such as a deviation of 3 points in the operational behavior mode, a deviation of 3 points in the state judgment indicator, and a fluctuation of 1 point in the interaction response characteristic (all within the allowable range); simultaneously determining the compensation requirements for each dimension from the rule set, such as compensating for the operational behavior mode deviation by 40%, compensating for the state judgment indicator by 60%, and requiring no compensation for the interaction response characteristic; then, starting the interpolation compensation algorithm to calibrate by dimension: the operational behavior mode is compensated with 3 × 40% = 1.2 points, resulting in a calibration score of 58 + 1.2 = 59.2 points; the state judgment indicator is compensated with 3 × 60% = 1.8 points, resulting in a calibration score of 52 + 1.8 = 53.8 points; the interaction response characteristic retains its initial score of 73 points; finally, calculating the overall calibration score according to preset weights (e.g., 40% for the operational behavior mode, 30% for the state judgment indicator, and 30% for the interaction response characteristic), integrating the scores of each dimension and the overall calibration score to form the calibrated scoring result.
[0060] Step 54: Based on the calibrated scoring results, a normalization method is used to standardize the scoring data to a standard numerical range, resulting in an optimized capability score. Specifically, this includes: first, determining the scoring standard range as 0 to 100 points based on industry standards; then, extracting the calibration scores for each dimension (operational behavior mode 59.2 points, state judgment index 53.8 points, and interaction response characteristics 73 points) and the overall calibration score (assuming 61.5 points) from the calibration results of Step 53; and statistically analyzing the range of calibration scores for each dimension (e.g., operation behavior mode 45 to 80 points, state judgment index 40 to 85 points, interaction response characteristics 55 points, etc.). (Up to 90 points); normalize the calibration scores for each dimension. For example, the operational behavior mode is calculated as (59.2-45) / (80-45)×100≈40.6 points, the status judgment index is calculated as (53.8-40) / (85-40)×100≈30.7 points, the interactive response characteristics are calculated as (73-55) / (90-55)×100≈51.4 points, and the overall calibration score is calculated as (61.5-40) / (90-40)×100=43 points. Check whether the normalization results are within 0 to 100 points. Finally, integrate the normalized scores of each dimension and the overall normalized scores to form the optimized capability score.
[0061] In this embodiment of the invention, a rule matching algorithm is used to match scoring data with adjustment rules, which can accurately associate the initial ability assessment score with the corresponding rules of each dimension in the personalized scoring adjustment rule set; avoid rule and data mismatch, and provide accurate matching basis for subsequent adjustments of each dimension; the deviation degree is calculated by the difference measurement algorithm, which can quantify the gap between the initial ability assessment score and the corresponding threshold of the rule, avoid subjective judgment bias, and the obtained scoring deviation data can clearly reflect the specific differences between the scores of each dimension and the rule requirements, providing data support for the calibration direction and magnitude; the initial score is calibrated by the interpolation compensation algorithm, which can smoothly adjust the initial score based on the scoring deviation data; avoid over-calibration or under-calibration, so that the calibrated result is more in line with the personalized adjustment rules, and improve the fit between the score and the rule; the normalization processing is used to standardize the score to the standard range, which can unify the data scale of the optimized ability score; eliminate the difference in the numerical range of scores of different dimensions, and make the score have a consistent reference standard.
[0062] In a preferred embodiment of the present invention, step 6 above involves multi-dimensional weighted aggregation of the optimized capability score to generate a comprehensive capability assessment result. This assessment result covers four virtual imaging dimensions: operational behavior patterns, interactive response characteristics, state judgment indicators, and execution stability, including: Step 61: Based on the optimized capability score, extract score data for four dimensions: operational behavior pattern, interactive response characteristics, state judgment indicators, and execution stability. Specifically, this includes: first, obtaining the optimized capability score, which includes normalized scores for the three dimensions of operational behavior pattern, interactive response characteristics, and state judgment indicators, as well as an overall normalized score; simultaneously, supplementing the execution stability dimension score data, which comes from the monitoring and analysis of dynamic indicators such as operational continuity, fault handling timeliness, and parameter control stability in power operations, and is obtained after standardization processing; then, extracting the normalized scores for the above three dimensions from the optimized capability score, such as 40.6 points, 51.4 points, and 30.7 points, and integrating the supplemented execution stability dimension score (such as 45.2 points) to form complete score data for the four dimensions.
[0063] Step 62: Based on the scoring data of the four dimensions, a weighted aggregation algorithm is used to calculate the comprehensive evaluation score through preset weight coefficients for each dimension. Specifically, this includes: First, determining the preset weight coefficients for the four dimensions according to the talent evaluation requirements of the power industry (operation behavior mode 35%, status judgment index 30%, execution stability 20%, interaction response characteristics 15%, total 100%); then, starting the weighted aggregation algorithm, multiplying the scoring data of each dimension by the corresponding weight coefficient to obtain the weighted score for each dimension; finally, adding the weighted scores of the four dimensions to calculate the comprehensive evaluation score.
[0064] Step 63: Based on the comprehensive evaluation score, the score is mapped to a preset evaluation level range through normalization processing to obtain a standardized evaluation result. Specifically, this includes: first, referring to the industry standard preset evaluation level range (Excellent 80 to 100 points, Good 60 to 79 points, Pass 40 to 59 points, Needs Improvement 0 to 39 points); then obtaining the comprehensive evaluation score from Step 62 (e.g., 42.3 points); mapping the comprehensive score to the corresponding level range through normalization processing, determining that 42.3 points falls within the Pass level range; and simultaneously recording the specific position of the comprehensive score within the level range to form a standardized evaluation result containing the comprehensive score, corresponding level, and position information within the level.
[0065] Step 64: Based on the standardized evaluation results, a comprehensive capability evaluation report covering four dimensions—operation behavior patterns, interactive response characteristics, state judgment indicators, and execution stability—is obtained. Specifically, this includes: first, organizing the four-dimensional scoring data from Step 61, the comprehensive evaluation score from Step 62, and the standardized evaluation results from Step 63; then analyzing the correlation between the scores of each dimension and the comprehensive result, for example, pointing out that a score lower than the comprehensive score requires special attention; next, writing the comprehensive capability evaluation report according to a fixed structure, beginning with an explanation of the evaluation object and cycle, presenting the scores of each dimension, capability analysis, comprehensive score and level in the middle, and providing targeted suggestions at the end; finally, checking the completeness of the report information and the accuracy of the data to form a comprehensive capability evaluation report covering the four virtual imaging dimensions.
[0066] In this embodiment of the invention, scoring data from four dimensions—operation behavior patterns, interaction response characteristics, state judgment indicators, and execution stability—is extracted. This clarifies the specific scoring sources for each dimension, providing complete and corresponding foundational data for subsequent multi-dimensional aggregation. It avoids incomplete data during aggregation due to missing dimensions, ensuring that each dimension participates in the comprehensive evaluation. A weighted aggregation algorithm is used to calculate the comprehensive evaluation score using preset dimension weight coefficients, allocating weights based on the differences in importance of each dimension in the evaluation. This avoids ignoring the differences in dimension importance through simple averaging, making the comprehensive score more aligned with evaluation needs and reflecting the different impacts of each dimension on overall capability. Standardization mapping the comprehensive score to a preset evaluation level range transforms abstract scores into clear level descriptions. It unifies the presentation standards of evaluation results, avoiding misunderstandings due to different score ranges, and facilitating a quick grasp of the comprehensive capability level. A comprehensive capability evaluation report covering all four dimensions is generated, integrating the scores of each dimension with the overall result. This avoids the one-sidedness of viewing dimension data or the comprehensive score in isolation, making the evaluation results more comprehensive and facilitating a holistic understanding of the performance of each dimension and the overall capability.
[0067] like Figure 2 As shown, embodiments of the present invention also provide a multimodal intelligent analysis system for assessing the capabilities of power industry personnel, comprising: The extraction module is used to perform weighted fusion on the multimodal input data that has been time-aligned and formatted according to a preset weight configuration strategy, and generate a fused multimodal data matrix; based on this matrix, multimodal feature vectors are extracted to represent operation behavior patterns, interaction response characteristics and state judgment indicators; The mapping module is used to input the multimodal feature vectors into a pre-trained neural network regression model for high-dimensional mapping calculation to obtain an initial capability assessment score. The matching module is used to extract typical modal features from speech and action modalities based on the initial ability assessment score, calculate the difference measure between them and the benchmark patterns in the virtual talent template library, and generate a feature difference vector; the feature difference vector is quickly matched and mapped using a geometric hash algorithm to obtain a suitable virtual type template. The rule set module is used to dynamically divide and merge regions in a three-dimensional feature space composed of operation behavior patterns, state judgment indicators and interaction response characteristics according to the virtual type template, and generate a scoring interval adapted to the current virtual imaging object; and to construct a personalized scoring adjustment rule set through interval mapping and normalization operations. The calibration module is used to calibrate the initial ability assessment score based on the personalized scoring adjustment rule set to obtain an optimized ability score; The evaluation results module is used to perform multi-dimensional weighted aggregation of the optimized capability scores to generate comprehensive capability evaluation results. The evaluation results cover four virtual imaging dimensions: operational behavior patterns, interactive response characteristics, state judgment indicators, and execution stability.
[0068] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0069] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0070] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A multimodal intelligent analysis method for assessing the capabilities of power industry personnel, characterized in that, The method includes: Step 1: Based on the preset weight configuration strategy, perform weighted fusion on the multimodal input data that has been time-aligned and formatted to generate a fused multimodal data matrix; extract multimodal feature vectors based on this matrix to characterize operation behavior patterns, interaction response characteristics and state judgment indicators; Step 2: Input the multimodal feature vector into the pre-trained neural network regression model for high-dimensional mapping calculation to obtain the initial capability assessment score; Step 3: Based on the initial ability assessment score, extract typical modal features from speech and action modalities, calculate the difference measure between them and the baseline patterns in the virtual talent template library, and generate feature difference vectors; use the geometric hash algorithm to quickly match and map the feature difference vectors to obtain the appropriate virtual type templates. Step 4: Based on the virtual type template, dynamically divide and merge regions in the three-dimensional feature space composed of operation behavior patterns, state judgment indicators and interactive response characteristics to generate a scoring interval adapted to the current virtual imaging object; construct a personalized scoring adjustment rule set through interval mapping and normalization operations. Step 5: Based on the personalized scoring adjustment rule set, calibrate the initial ability assessment score to obtain the optimized ability score; Step 6: Perform multi-dimensional weighted aggregation on the optimized capability score to generate a comprehensive capability assessment result. The assessment result covers four virtual imaging dimensions: operational behavior pattern, interactive response characteristics, state judgment index, and execution stability.
2. The multimodal intelligent analysis method for assessing the capabilities of power industry personnel according to claim 1, characterized in that, Step 1 includes the following: Collect multimodal raw data streams generated during power operation and maintenance, including text records, voice signals, motion sensor sequences, and numerical logs; preprocess the multimodal raw data streams to obtain time-aligned and formatted multimodal data; Based on the regularized multimodal data, the corresponding work environment level and operation task type are identified; according to the work environment level and operation task type, a fusion weight coefficient is dynamically configured for each modal data to generate a weight allocation scheme.
3. The multimodal intelligent analysis method for assessing the capabilities of power industry personnel according to claim 2, characterized in that, Step 1: Based on the preset weight configuration strategy, perform weighted fusion on the multimodal input data that has been time-aligned and formatted, and generate a fused multimodal data matrix; Based on this matrix, multimodal feature vectors are extracted to characterize operational behavior patterns, interaction response characteristics, and state judgment indicators, including: Based on the weight allocation scheme, the regularized multimodal data is aligned in the temporal and spatial domains to generate a spatiotemporally aligned multimodal data sequence. Based on the spatiotemporally aligned multimodal data sequence, a weighted fusion calculation is performed according to the fusion weight coefficients corresponding to each modality to generate a weighted fused multimodal data matrix; Principal component analysis was performed on the weighted and fused multimodal data matrix to obtain the dimensionality-reduced principal component feature set; Based on the principal component feature set, the final discriminant projection vector is obtained through linear discriminant analysis; The principal component feature set is spatially projected using the final discriminant projection vector to extract the feature combination with the maximum class discrimination, forming a multimodal feature vector for characterizing operational behavior patterns, interactive response characteristics, and state judgment indicators.
4. The multimodal intelligent analysis method for assessing the capabilities of power industry personnel according to claim 3, characterized in that, Step 2 involves inputting the multimodal feature vectors into a pre-trained neural network regression model for high-dimensional mapping calculation to obtain an initial capability assessment score, including: Based on the multimodal feature vectors, the Z-score normalization method is used to perform feature scaling to obtain the normalized feature vectors. The standardized feature vector is input into a pre-trained neural network regression model; forward propagation calculation is performed through the input layer, hidden layer and output layer of the neural network regression model, and a multi-dimensional initial capability evaluation vector is obtained from the output layer. Based on the initial capability assessment vector, a nonlinear transformation process is performed using the Sigmoid activation function to obtain the transformed vector data. Based on the transformed vector data, it is mapped to a preset standard numerical range through linear scaling calculation to obtain a standardized output vector; Based on the standardized output vector and the preset weight coefficients corresponding to the operation behavior mode, interaction response characteristics and state judgment indicators, a weighted summation algorithm is used to calculate the weighted sum, which is the initial capability assessment score.
5. The multimodal intelligent analysis method for assessing the capabilities of power industry personnel according to claim 4, characterized in that, Step 3: Based on the initial ability assessment score, extract typical modal features from the speech and action modalities, calculate the difference measure between them and the benchmark patterns in the virtual talent template library, and generate a feature difference vector, including: Based on the initial capability assessment score, a subset of key features is selected from the raw feature data of speech and action modalities; Based on a subset of key features, a distance metric algorithm is used to calculate the degree of difference between it and each benchmark pattern in the virtual talent template library, and the initial difference metric results are obtained. Based on the initial difference measurement results, dimensionality reduction was performed using principal component analysis to obtain the dimensionality-reduced difference features. Based on the differences after dimensionality reduction, cluster analysis is used to classify patterns and identify the closest baseline pattern category. Based on the benchmark pattern category, the corresponding benchmark feature vector is extracted, and the difference between the key feature subset and the benchmark feature vector in each dimension is calculated; Based on the differences in each dimension, a standardized feature difference vector is generated through normalization.
6. The multimodal intelligent analysis method for assessing the capabilities of power industry personnel according to claim 5, characterized in that, The geometric hash algorithm is used to quickly match and map the feature difference vectors to obtain a suitable virtual type template, including: Based on the aforementioned feature difference vector, dimensionality reduction is performed using principal component analysis to obtain low-dimensional features. Based on the aforementioned low-dimensional features, cluster analysis is used to classify patterns and identify their similarity to each benchmark pattern in the virtual talent template library. Based on the similarity, the top K benchmark patterns with the highest similarity are selected as the candidate template set; Based on the candidate template set, the corresponding baseline feature vectors are extracted, and a hash table structure is constructed. Based on the hash table structure, a geometric hash algorithm is used for fast matching to calculate the matching degree between the feature difference vector and each candidate template; Based on the matching degree, the candidate template with the highest matching degree is selected as the adapted virtual type template.
7. The multimodal intelligent analysis method for assessing the capabilities of power industry personnel according to claim 6, characterized in that, Step 4: Based on the virtual type template, dynamically divide and merge regions in the three-dimensional feature space composed of operation behavior patterns, state judgment indicators and interactive response characteristics to generate a scoring range adapted to the current virtual imaging object. Through interval mapping and normalization operations, a personalized score adjustment rule set is constructed, including: Based on the virtual type template, a three-dimensional feature space is constructed, consisting of three dimensions: operation behavior mode, state judgment index, and interaction response characteristics. Based on the three-dimensional feature space, cluster analysis is used to divide the initial region, resulting in multiple initial feature regions. Based on the initial feature region, and using the feature distribution pattern defined in the virtual type template as the merging basis, a merging operation is performed on regions with similar and adjacent feature distributions in the space to obtain an optimized feature region division. Based on the optimized feature region division, the feature distribution parameters of each region are extracted to obtain a dynamic scoring range adapted to the current virtual imaging object. Based on the dynamic scoring interval, a linear mapping algorithm is used to map the original scoring data to the corresponding interval range to obtain the interval mapping result; Based on the interval mapping results, data standardization is performed using a normalization method to form a set of personalized scoring adjustment rules.
8. The multimodal intelligent analysis method for assessing the capabilities of power industry personnel according to claim 7, characterized in that, Step 5: Based on the personalized scoring adjustment rule set, calibrate the initial ability assessment score to obtain the optimized ability score, including: Based on the personalized scoring adjustment rule set and the initial ability assessment score, a rule matching algorithm is used to perform matching calculations between the scoring data and the adjustment rules to obtain the rule matching results for each dimension. Based on the rule matching results, a difference measurement algorithm is used to calculate the degree of deviation between the initial ability assessment score and the corresponding threshold of the personalized score adjustment rule set, and score deviation data is obtained. Based on the scoring deviation data, the initial ability assessment score is calibrated using an interpolation compensation algorithm to obtain the calibrated scoring result; Based on the calibrated scoring results, a normalization method is used to standardize the scoring data to a standard numerical range, resulting in an optimized capability score.
9. The multimodal intelligent analysis method for assessing the capabilities of power industry personnel according to claim 8, characterized in that, Step 6: Perform multi-dimensional weighted aggregation on the optimized capability scores to generate a comprehensive capability assessment result. This assessment result covers four virtual imaging dimensions: operational behavior patterns, interactive response characteristics, state judgment indicators, and execution stability, including: Based on the optimized capability score, score data for four dimensions are extracted: operational behavior pattern, interaction response characteristics, state judgment index, and execution stability. Based on the scoring data of the four dimensions, a weighted aggregation algorithm is used to calculate the comprehensive evaluation score through preset weight coefficients for each dimension. Based on the comprehensive evaluation score, the score is mapped to a preset evaluation level range through normalization processing to obtain a standardized evaluation result; Based on the standardized evaluation results, a comprehensive capability evaluation report covering four dimensions—operational behavior patterns, interactive response characteristics, status judgment indicators, and execution stability—is obtained.
10. A multimodal intelligent analysis system for assessing the capabilities of power industry personnel, wherein the system implements the method as described in any one of claims 1 to 9, characterized in that, include: The extraction module is used to perform weighted fusion on multimodal input data that has been time-aligned and formatted based on a preset weight configuration strategy, and generate a fused multimodal data matrix. Based on this matrix, multimodal feature vectors are extracted to characterize operational behavior patterns, interactive response characteristics, and state judgment indicators. The mapping module is used to input the multimodal feature vectors into a pre-trained neural network regression model for high-dimensional mapping calculation to obtain an initial capability assessment score. The matching module is used to extract typical modal features from speech and action modalities based on the initial ability assessment score, calculate the difference measure between them and the benchmark patterns in the virtual talent template library, and generate a feature difference vector. The geometric hash algorithm is used to quickly match and map the feature difference vectors to obtain a suitable virtual type template. The rule set module is used to dynamically divide and merge regions in a three-dimensional feature space composed of operation behavior patterns, state judgment indicators and interaction response characteristics according to the virtual type template, and generate a scoring range that is adapted to the current virtual imaging object. A set of personalized score adjustment rules is constructed through interval mapping and normalization operations; The calibration module is used to calibrate the initial ability assessment score based on the personalized scoring adjustment rule set to obtain an optimized ability score; The evaluation results module is used to perform multi-dimensional weighted aggregation of the optimized capability scores to generate comprehensive capability evaluation results. The evaluation results cover four virtual imaging dimensions: operational behavior patterns, interactive response characteristics, state judgment indicators, and execution stability.