An event extraction and prediction method based on temporal events and semantic context

By combining a temporal event- and semantic context-based approach with the NILA-GCN model and random forest algorithm, the problems of inaccurate time prediction and insufficient user intent recognition in existing technologies are solved, achieving high-precision event prediction and improved user experience.

CN114241381BActive Publication Date: 2025-10-31GUANGDONG OPEN UNIV (GUANGDONG POLYTECHNIC VOCATIONAL COLLEGE)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111548478.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-10-31
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

Existing technologies lack standardized time prediction methods for event prediction, resulting in a mixed time stream and an inability to effectively extract time-series events. Furthermore, rule-based and statistical methods cannot accurately identify user intent, leading to limitations in dialogue systems and a poor user experience.

Method used

We employ an event extraction and prediction method based on temporal events and semantic context. By acquiring images and videos and collecting semantic context, we construct the NILA-GCN model and the semantic context model. Combined with the random forest algorithm and the n-grams model, we achieve data feature extraction and prediction.

Benefits of technology

It improves the accuracy of event prediction and user experience, and can instantly acquire video data information, convert it into time series and audio background, construct a time series event identification domain, and achieve high-precision timestamp recording and semantic information processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114241381B_ABST
    Figure CN114241381B_ABST
Patent Text Reader

Abstract

This invention discloses a method for event extraction and prediction based on temporal events and semantic context, comprising the following steps: (S1) Real-time acquisition of event data information, using image acquisition, video acquisition, and semantic context acquisition methods; (S2) Storage of the acquired video stream to obtain video stream data information; simultaneously generating a temporal event stream from the acquired video stream data information and recording timestamps; simultaneously converting the acquired data information into short text samples and recording short text sample labels; (S3) Extraction of data features, using an image recognition model to extract data stream information features from the video stream data information, and using a constructed classifier to analyze the short text samples; (S4) Construction of a prediction model for the predicted event and a semantic context model to achieve data prediction; (S5) Outputting the prediction results. This invention enables the extraction and prediction of events based on temporal events and semantic context, improving event prediction capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of time prediction and evaluation, and more specifically to a method for event extraction and prediction based on time-series events and semantic context. Background Technology

[0002] Event prediction has become a hot and challenging area of ​​application in the field of computer vision in recent years. With the rapid development of computer, storage, and network technologies, and the continuous updates to various digital devices and mobile terminals, the volume of data is growing explosively. Existing technologies suffer from the following technical shortcomings:

[0003] (1) There is a lack of standards for time prediction. When analyzing event information, the time flow is mixed and messy, making it impossible to extract time-series events from existing event data information. The extraction accuracy is low, which leads to poor time assessment and prediction capabilities.

[0004] (2) Lack of rules, rule-based methods, and statistical methods. Rule-based methods refer to using computer language to describe rules based on defined grammar rules, parts of speech, word formation, and sentence construction rules. Statistical methods refer to using deep learning and big data to build dialogue systems and automatically generate dialogues. In practical use, existing dialogue recognition systems have low ability to identify user intent. They often fail to answer user questions due to the inability to determine intent, or provide irrelevant or repetitive answers, resulting in limited dialogue content and a poor user experience. Summary of the Invention

[0005] To address the shortcomings of the aforementioned technologies, this invention discloses an event extraction and prediction method based on temporal events and semantic context, which can realize the extraction and prediction of events based on temporal events and semantic context, thereby improving the event prediction capability.

[0006] The present invention adopts the following technical solution:

[0007] An event extraction and prediction method based on temporal events and semantic context includes the following steps:

[0008] (S1) Real-time acquisition of event data information, through image acquisition, video acquisition and semantic background acquisition;

[0009] (S2) Store the acquired video stream to obtain video stream data information; at the same time, generate a time-series event stream from the acquired video stream data information and record the timestamp; at the same time, convert the acquired data information into short text samples and record the short text sample labels;

[0010] (S3) Extract data features: Video stream data information is extracted using an image recognition model, and short text samples are analyzed using a constructed classifier.

[0011] (S4) Construct a predictive event model and a semantic context model to achieve data prediction;

[0012] (S5) Output the prediction results.

[0013] As a further technical solution of the present invention, the real-time acquisition of event data information in step (S1) is based on an embedded multi-channel digital video acquisition device, including a video input interface, a video acquisition module, a core processor module, a video storage module, an external interface module, and a video transmission module. The output of the video input interface is connected to the input of the video acquisition module, the output of the video acquisition module is connected to the input of the core processor module, the output of the core processor module is connected to the input of the video storage module, the output of the core processor module is also connected to the input of the external interface module, and the video storage module is also connected to the video transmission module. The video main control chip uses either a TMS320DM8168 chip or a TVP5158 chip, and the main control chip includes an ARM module, a video processing module, an OCR recognition module, and a DSP module; the TVP5158 chip includes an FPGA module. The above technical solution constitutes an overall solution for real-time acquisition of event data information, enabling real-time acquisition of event data information.

[0014] As a further technical solution of the present invention, in step (S2):

[0015] The method for generating a time-series event stream from the acquired video stream data and recording timestamps is as follows: An mVideo data acquisition interface is set, the data stream input time and transmission time are set, a time-series event identification domain is constructed, and data falling within the time-series event identification domain is recorded as a timestamp. The method for recording short text sample labels is a text similarity evaluation method, which employs an n-grams model-based text-to-text similarity method.

[0016] As a further technical solution of the present invention, the method for data information analysis in step (S3) is the random forest algorithm.

[0017] As a further technical solution of the present invention, in step (S4), the method for constructing the prediction event prediction model is the NILA-GCN model, which includes an input layer, a convolutional layer, a fusion layer, and a loss function module, wherein the output end of the input layer is connected to the input end of the convolutional layer, the output end of the convolutional layer is connected to the input end of the fusion layer, and the output end of the fusion layer is connected to the input end of the loss function module.

[0018] As a further technical solution of the present invention, in step (S4), the prediction method of the prediction event prediction model is as follows:

[0019] (S41) Population initialization: Input 2n video data parameters and distribute them equally to two groups as candidate populations, namely male lionAm=[y1, y2, y3, ···, yi] and female lionAf=[y1, y2, y3, ···, yi];

[0020] (S42) Crossover and mutation;

[0021] The acquisition of video data parameters generates new individuals through crossover and mutation. A crossover algorithm based on dual probabilities is used to realize the probabilities of two different data types. Male lionAm and female lionAf mate to generate a new acquisition of video data parameter A. cub =[y1, y2, y3,···,yi];

[0022] (S43) Territorial defense: During the process of iteratively generating superior individuals from the acquired video data parameters, male lionAm and female lionAf will be attacked by external abnormal information lions; at this time, male lionAm will defend and protect the superior acquired video data parameters and delineate the area around the population as territory.

[0023] (S44) The optimal solution is obtained. The inferior solution between male lionAm and female lionAf will be replaced with the optimal solution. The crossover will not end before the iteration termination condition is reached. After the iteration termination condition is reached, the crossover ends and the optimal solution is output.

[0024] As a further technical solution of the present invention, in step (S4), the semantic background model construction model includes a database, an analyzer, an n-grams model, a clustering model, and a comprehensive evaluation model, wherein the output of the database is connected to the input of the analyzer, the output of the analyzer is connected to the input of the n-grams model, the output of the n-grams model is connected to the input of the clustering model, and the output of the clustering model is connected to the input of the comprehensive evaluation model.

[0025] As a further technical solution of the present invention, in step (S4), the prediction method of the semantic background model is as follows:

[0026] (1) The system server collects semantic parameter topic data by outputting data information through the database; the system server submits the user's query terms to the search engine and allows the user to select the preferred semantics in the results of the returned page and form a new semantic set for the user;

[0027] (2) Semantic information is obtained through the analyzer, and the system server establishes a new semantic model for the user; the system server establishes a concept graph reflecting the user's preferences through the new semantic set; the system server first constructs a concept lattice before establishing the new semantic model for the user;

[0028] (3) A new semantic background is constructed through the n-grams model, and the system server establishes a conceptual semantic background graph; the system server converts the concept lattice into a conceptual semantic background graph that can intuitively represent the semantic relationships between semantics;

[0029] (4) The acquired speech information is analyzed and calculated in different forms through clustering model and comprehensive evaluation model. The system server updates the concept semantic background map to update the new semantic data of the user. The system server increases or decreases the concept semantic background map and finally outputs semantic information with high purity.

[0030] As a further technical solution of the present invention, the method for recording the timestamp in step (S2) is as follows:

[0031] Configure the mVideo data acquisition interface, set the data stream input time and transmission time, construct a time-series event identification domain, and record data information falling within the time-series event identification domain as a timestamp. The time-series event identification domain construction process involves comparing historical time-series event data with the model, then calculating the difference between the historical data and the model data. If more than 96% of the perceived video data information points fall within the area surrounding the model, then the formula is activated.

[0032]

[0033] When formula (1) is a polynomial equation that approximates the real time sequence event, the principle of formula (1) is the principle of linear algebra, which connects different data information through linear connection to represent the data information of the perceived video data information points.

[0034] As a further technical solution of the present invention, the text similarity method based on the n-grams model is as follows:

[0035] First, all semantic texts related to the event are retrieved from the text database. Then, appropriate preprocessing is performed to obtain an n-grams semantic model. The numerical expression or frequency matrix of the text is found for visualization, clustering, and numerical estimation. After creating the frequency matrix, the matrix is ​​provided to the SOM. The dataset is clustered and visualized in the SOM to detect the visual similarity of the texts. At the same time, four algorithms, namely the cosine similarity method, the dice similarity method, the extended Jaccard similarity method, and the overlap coefficient similarity method, are used to calculate the similarity measure or numerical estimate.

[0036] In the process of creating the frequency matrix, the final calculation result depends on the creation of n-gram words, including the selected filters. n-gram words include sentences, paragraphs, keywords, and semantics.

[0037] Positive and beneficial effects:

[0038] This invention can acquire video data information in real time and convert event video data information into time sequence and audio background. By multiplexing the acquired data information, it can simultaneously convert four analog video signals into digital video signals. It can receive digital video signals in two formats: ITU-R BT.656 and BT.1120. Event video stream recording is achieved by recording timestamps in the video data information.

[0039] This invention can also construct a time-series event identification domain, compare historical data of time-series events with the model, and then calculate the difference between the historical data and the model data to record the timestamp.

[0040] This invention presents a text similarity method based on an n-grams model, estimating results from both visual and numerical perspectives. Semantic information is processed through text preprocessing, visual clustering, and numerical estimation. During semantic evaluation, a frequency matrix is ​​also created to ultimately calculate the evaluation results.

[0041] This invention also employs the random forest algorithm to classify and process semantic information, and uses the NILA-GCN model to construct a spatiotemporal convolutional structure to improve semantic analysis capabilities. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:

[0043] Figure 1 This is a schematic diagram of the overall architecture for event evaluation in this invention;

[0044] Figure 2 This is a schematic diagram of the event video acquisition module of the present invention;

[0045] Figure 3 This is a schematic diagram of the video acquisition circuit in the event video acquisition module of the present invention;

[0046] Figure 4 This is a schematic diagram of the timing discrimination domain of the present invention;

[0047] Figure 5 This is a schematic diagram of the text construction model of the present invention;

[0048] Figure 6 This is a schematic diagram illustrating the construction of the random forest algorithm model of this invention;

[0049] Figure 7 This is a classification diagram illustrating one embodiment of the random forest algorithm of the present invention;

[0050] Figure 8 This is a schematic diagram of the NILA-GCN model architecture of the present invention;

[0051] Figure 9 This is a schematic diagram of the spatiotemporal convolution structure of the convolution module of the present invention;

[0052] Figure 10 This is a schematic diagram of the prediction method of the prediction model in the prediction event of the present invention. Detailed Implementation

[0053] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0054] like Figure 1 As shown, an event extraction and prediction method based on temporal events and semantic context includes the following steps:

[0055] (S1) Real-time acquisition of event data information, through image acquisition, video acquisition and semantic background acquisition;

[0056] (S2) Store the acquired video stream to obtain video stream data information; at the same time, generate a time-series event stream from the acquired video stream data information and record the timestamp; at the same time, convert the acquired data information into short text samples and record the short text sample labels;

[0057] (S3) Extract data features: Video stream data information is extracted using an image recognition model, and short text samples are analyzed using a constructed classifier.

[0058] (S4) Construct a predictive event model and a semantic context model to achieve data prediction;

[0059] (S5) Output the prediction results.

[0060] To make the embodiments of the present invention clearer, the embodiments of the present invention will be described in detail below.

[0061] like Figure 2 As shown, the real-time acquisition of event data in step (S1) is based on an embedded multi-channel digital video acquisition device, including a video input interface, a video acquisition module, a core processor module, a video storage module, an external interface module, and a video transmission module. The output of the video input interface is connected to the input of the video acquisition module, the output of the video acquisition module is connected to the input of the core processor module, the output of the core processor module is connected to the input of the video storage module, the output of the core processor module is also connected to the input of the external interface module, and the video storage module is also connected to the video transmission module. The main video control chip uses either a TMS320DM8168 chip or a TVP5158 chip, and the main control chip includes an ARM module, a video processing module, an OCR recognition module, and a DSP module; the TVP5158 chip includes an FPGA module. This technical solution realizes the acquisition, storage, interaction, and processing of data information.

[0062] In the above embodiments, the main control chip of the multi-channel digital video acquisition device adopts the TMS320DM8168 chip, which supports local hard disk storage, local real-time display, and playback of video data, and acts as the core control module to coordinate the control of other modules. The main control chip integrates multiple processing cores, including an ARM subsystem, a high-definition video processing subsystem, an encoding / decoding subsystem, and a DSP subsystem. The ARM subsystem is responsible for configuring and controlling other peripheral circuits, the high-definition video processing subsystem is responsible for compression encoding, filtering, and format conversion of video data, and the DSP subsystem is responsible for video data management. The video acquisition module uses the TVP5158 chip, which is responsible for acquiring multiple video streams from the hydropower station and sending the acquired video data to the main control module for processing. By adopting the TMS320DM8168 chip, control of video data acquisition can be achieved. The TMS320DM8168 is a high-performance video processor combining a floating-point DSP C674x and an ARM Cortex-A8. The main frequency parameters are 930MHz (DSP) + 1.1GHz (ARM). In one specific embodiment, the system controls the image sensor via a serial port, allowing three-channel image data signals, a clock, and various synchronization signals to be input as required. The system then sequentially acquires, processes, and stores the image signals. The system utilizes its built-in interface to achieve display, host computer communication, keyboard control, and other functions, enabling user-friendly human-machine interaction. This system uses the latest TMS320DM8168 chip from TI's DaVinci series. This chip integrates a 1GHz ARM Cortex-A8, a 1GHz TI C674x floating-point DSP, several second-generation programmable high-definition video image coprocessors, an innovative high-definition video processing subsystem (HDVPSS), and a comprehensive codec. It supports high-definition resolutions including H.264, MPEG-4, and VC1, and includes multiple interfaces such as Gigabit Ethernet, PCI Express, SATA2, DDR2, DDR3, USB2.0, MMC / SD, HDMI, and DVI, supporting more functional expansion and complex applications. This chip was used to design and implement the acquisition, processing, and display of two or three image signals with different resolutions. The hardware schematic is shown below. Figure 2 As shown. The hardware modules involved in the development and design of this system include: an image acquisition interface module, an image acquisition module, an image storage module, and a peripheral interface module. During image data acquisition, the TMS320DM8168's HDVPSS (HDVideo Processing Subsystem) provides video input and output interfaces. The video input interface allows access to external image devices (such as image sensors, video decoders, etc.).

[0063] Video stream data information can be any kind of information, such as hydropower station operation data, industrial site monitoring data, on-site operation data of actors, and other types of time-based video data streams.

[0064] like Figure 3 As shown in the above embodiment, the video acquisition module employs multi-channel video multiplexing to simultaneously convert four analog video signals into digital video signals. The acquired hydropower station video data first needs to be processed by an AD converter, then the chroma and luminance signals of the video image are separated, filtered, and integrated. The video signal is then processed by a scaler to determine its format, and finally, according to the system's requirements for the hydropower station video images, the digital video format is sent to the input interface of the back-end video processor. The video acquisition module circuit supports 16 channels of digital video input and can receive digital video signals in both ITU-R BT.656 and BT.1120 formats. The multiple video signals input to the decoder input of the TVP5158 chip are digitized and multiplexed into a single output, improving the utilization rate of the main control chip's VP port.

[0065] In one specific embodiment, the High-Definition Video Processing Subsystem (HDVPSS) has two independent video capture input ports, VIP0 and VIP1. VIP0 can be configured in 24-bit, 16-bit, and two independent 8-bit modes, while VIP1 can be configured in 16-bit and two independent 8-bit modes. From the capture frequency and various configuration modes, it can be seen that multiple implementation methods are possible for different traffic volumes. For simplicity in storage design, this solution configures VIP0 for 24-bit acquisition. In this mode, the maximum traffic volume is 165MB × 248 = 495MB / s, which meets the traffic requirements. After data acquisition, the HDVPSS subsystem is configured to send the data to VPDMA, and finally to DDR memory. When the data volume in DDR memory reaches the set data volume, an interrupt is generated. After the interrupt occurs, DMA transfer between memory and solid-state drive is initiated according to the storage address, storing the acquired image on the SSD through the SATA2 interface, thus achieving data storage.

[0066] like Figure 4 As shown, the method for generating a time-series event stream from the acquired video stream data and recording timestamps is as follows: Set the mVideo data acquisition interface, set the data stream input time and transmission time, construct a time-series event identification domain, and record data information falling within the time-series event identification domain as a timestamp. The time-series event identification domain construction process involves comparing historical time-series event data with the model, then calculating the difference between the historical data and the model data. If more than 96% of the perceived video data information points fall within the area surrounding the model, then the formula is activated:

[0067]

[0068] When formula (1) is a polynomial equation approximating a real temporal event, the principle of formula (1) is the principle of linear algebra, which connects different data information through linear connections to represent the data information of perceived video data points, such as different data points A, B, C, etc., and connects them together. k a represents a data series of any point. k This represents the coefficient value of the collected data. There is inevitably an error between the time-series event model and the actual discrete sensing data points. This error is the distance between each discrete point in the extracted video features and the standard model, and can be represented by an error function:

[0069]

[0070] Where a represents the coefficient vector of the regression equation, x n Let y be the x-coordinate of the nth point in the dataset, where y n Represented as x n The ordinate value, (x n ,y n This refers to actual perceived data. For example... Figure 4 As shown, points falling within a certain range on the line are assigned a timestamp, while points falling outside the line are removed.

[0071] like Figure 5 As shown, the method for recording short text sample labels in step (S2) is a text similarity evaluation method, which adopts a text similarity method based on the n-grams model.

[0072] First, all semantic text related to the event is retrieved from a text database, and then appropriate preprocessing is performed to obtain an n-gram semantic model. The main purpose of text preprocessing is to find the numerical expression of the text (finding the frequency matrix) for visualization, clustering, and numerical estimation. After creating the frequency matrix, the matrix is ​​provided to the SOM (Signal-Oriented Model), where the dataset is clustered and visualized to detect the visual similarity of the texts. Simultaneously, four algorithms—cosine similarity, dice similarity, extended Jaccard similarity, and overlap coefficient similarity—are used to calculate similarity measures (numerical estimation). The SOM helps to visualize the similarity of the entire text dataset on a map, while the numerical estimation framework can quantitatively demonstrate and specify the results. Combining these two techniques allows for deeper text similarity analysis, ultimately leading to a comprehensive evaluation result.

[0073] In the process of creating the frequency matrix, the final calculation result depends on the filter selected when creating the n-grams word package. Therefore, choosing an appropriate filter to obtain accurate results is crucial. The n-grams package is a semantic probabilistic model defined on a sequence of words. It can analyze all text related to the cause of a power outage, or simply divide it into several parts: sentences, paragraphs, keywords, and semantics. Depending on the text, the n-grams model is formed at the character level or word level.

[0074] like Figure 6 As shown, the event data classification algorithm designed in this invention can be divided into three modules. The data acquisition module collects relevant data information, monitors the relevant data information through the Flume component, and stores and manages the data through queue information and a distributed data management system. The data processing module performs data packet clustering on the stored data and extracts network data features and processes them in the form of data streams. Finally, the improved and optimized random forest algorithm is used to train the model data and generate the required data classification model.

[0075] like Figure 7 As shown, in order to establish a decision tree similarity matrix, this invention needs to transform the tree-structured decision tree into a binary tree structure. Each sub-scheme of the transformed binary tree decision tree can be transformed into a set similar to {"A<=X", "B>Y", "C<=Z", "class_3"}. By comparing the elements in the sub-scheme set, the similarity between the two schemes can be obtained.

[0076] To compare the similarity of each element in the set of sub-solutions, the present invention provides a formula for calculating the similarity Simn of non-numerical elements, as shown in Equation 1.

[0077]

[0078] The formula for calculating the element similarity Simn for numerical elements is shown in Equation 4.

[0079]

[0080] As shown in Equation 4, k1 and k2 are numerical values, representing the upper and lower limits of the vertical discriminant, and t represents the maximum similarity that the user can tolerate between the two schemes.

[0081] The similarity Sim of the decision tree sub-schemes can be calculated by averaging the similarity of each element, as shown in Equation 5.

[0082]

[0083] The similarity equation between the two decision trees can be calculated by calculating the similarity of the sub-schemes, as shown in Equation 4.

[0084]

[0085] The similarity between any two decision trees can be established by calculating the similarity of all decision trees in the random forest algorithm, as shown in Equation 7.

[0086]

[0087] As shown in Equation 7, by observing the similarity matrix of the random forest algorithm, we can find that the matrix Sim n,m The similarity between the nth and mth decision trees is represented by the similarity matrix. Analyzing the similarity matrix of the algorithm can select the better result to achieve the integration of decision trees.

[0088] like Figure 8 As shown, in order to accurately predict events, this invention designs the overall architecture of the GCN model based on the time dimension. To obtain sufficient time dimension information, this invention describes the recent parameters (x) of the event occurrence. h ), daily cycle parameter (x) d ) and periodic parameter (x) w The GCN model, featuring three temporal parameters, consists of multiple graph convolutional modules and a fully connected (FC) layer sharing the same neural network structure. The graph convolutional module design includes both spatial graph convolution operations and standard 2D convolution operations in the temporal dimension.

[0089] Because different nodes are affected to varying degrees by the characteristics of parameters in different time dimensions, the GCN model ultimately fuses the three outputs based on the parameter matrix to fully leverage the role of multiple components. The fused result is the final prediction. The prediction method can be specifically...

[0090] Step 1: Population Initialization

[0091] In the first stage of this LA algorithm, 2n event data parameters are evenly distributed to two groups as candidate populations, namely male lionA. m =[y1, y2, y3, ..., y i ] and female lionA f =[y1, y2, y3, ..., y i ]. Where i represents the length of the population solution vector.

[0092] Step 2: Crossover and Mutation

[0093] In the LA algorithm, new individuals are generated from event data parameters through crossover and mutation. This invention proposes a crossover algorithm based on two probabilities, that is, using two different probabilities to achieve crossover. Male lionA m and female lionA f New event data parameter A was generated through mating. cub =[y1, y2, y3, ..., y i Among them, four new event data parameters A 1~4 Two randomly selected intersection points are used to generate four new event data parameters, intersection point A, through p-random mutation. 5~8 After crossover and mutation, a total of 8 new event data parameters were generated, which were then divided into male lionA groups using K-means clustering. m and female lionA f Then, based on the quality of different event data parameters, the worst event data parameters are filtered out, and the male lionA is ensured to be eliminated. m and female lionA f The number of event data parameters is equal. Iterative calculation and re-initialization of the population.

[0094] Step 3: Territory Defense

[0095] During the iterative generation of superior individuals from event data parameters, male lionA m and female lionA f They will be attacked by lions from outside the population. At this time, male lion A... m The data parameters of the excellent event will be protected and safeguarded, and the area around the population will be designated as territory.

[0096] The generation method of individual lions and the territory of male lions A m and female lionA f The same, then use the new solution y n Attack the population of event data parameters within the attack territory. If the individual lion's solution y n If a solution is preferred over other solutions in this process, then y is used. n Replace the event data parameter y of the population within the territory i The new lion will continue the crossover and mutation process. The objective function f(y) of the LA algorithm can be calculated according to equation (8):

[0097]

[0098] In equation (8), f(y) m ) and f(y f () represent male lion A m and female lionA fThe objective function value, f(y) m_cub ) and f(y f_cub ) represent the objective function values ​​of male and female lions in the new event data, respectively. m_cub || refers to the number of male lions in the new event data parameters.

[0099] Step 4: Obtain the optimal solution

[0100] In this step, male lionA m and female lionA f Inferior solutions will be replaced with optimal solutions, and the crossover will not end until the iteration termination condition is met. The optimal solution y of the LA algorithm... b Determined according to the following inequality (3):

[0101] f(y b )<f(y),y b ≠y (9)

[0102] In the iterative calculation of the LA algorithm, there exists a parameter g representing the number of reproductions, g b This refers to the optimal reproductive capacity of the event data parameter population, typically set to 5. `g` is set to 0 in the first generation and gradually increases. If the female lion event data parameter population is replaced, `g` must start from 0. After completing the above steps, return to step 2 until the iteration termination condition is met, obtaining the optimal solution of the LA algorithm.

[0103] NI's optimized LA algorithm can perform the following operations:

[0104] Based on the value of the objective function, within the specified iteration interval, obtain the cloning event data parameter M at the location center of the lion population:

[0105]

[0106] In equation (10), M j M is the clone number of the j-th event data parameter. max This indicates that the maximum number of clones is set to 40 here, P j is the objective function value of the j-th event data parameter, and N refers to the number of lions in the population. After cloning, individual lions mutate the M cloned event data parameters. Mutations are performed on lion populations with lower objective function values, as shown in equations (9) and (10).

[0107] x i+1 =x i +r×randn(1) (10)

[0108]

[0109] In equations (10) to (11), x refers to the lion population, x i+1 These are the new event data parameters after same-sex crossover, where r refers to the radius of other lion populations from the center of the mutant population; P max This represents the maximum value at the center of the lion population. The M-mutated lion population is screened and compared, and the lion with the largest objective function value is selected. The optimal solution for the LA algorithm is then calculated.

[0110] Figure 9 This is a schematic diagram of the spatiotemporal convolution structure of the graph convolution module of the present invention; the GCN model finally fuses the three output results based on the parameter matrix to give full play to the role of multiple components, and the fused result is the final prediction result.

[0111] like Figure 10 As shown, the prediction method of the semantic background model is as follows:

[0112] The semantic background model is constructed by a database, an analyzer, an n-grams model, a clustering model, and a comprehensive evaluation model. The output of the database is connected to the input of the analyzer, the output of the analyzer is connected to the input of the n-grams model, the output of the n-grams model is connected to the input of the clustering model, and the output of the clustering model is connected to the input of the comprehensive evaluation model.

[0113] (1) The system server collects semantic parameter topic data; the system server submits the user's query terms to the search engine and allows the user to select the preferred semantics in the results of the returned page, thus forming a new semantic set for the user;

[0114] (2) Semantic information is obtained through the analyzer, and the system server establishes a new semantic model for the user; the system server establishes a concept graph reflecting the user's preferences through the new semantic set; the system server first constructs a concept lattice before establishing the new semantic model for the user;

[0115] (3) A new semantic background is constructed through the n-grams model, and the system server establishes a conceptual semantic background graph; the system server converts the concept lattice into a conceptual semantic background graph that can intuitively represent the semantic relationships between semantics;

[0116] (4) The acquired speech information is analyzed and calculated in different forms through clustering model and comprehensive evaluation model. The system server updates the concept semantic background map to update the new semantic data of the user. The system server increases or decreases the concept semantic background map and finally outputs semantic information with high purity.

[0117] The following is a simplified implementation example. Assume a provincial power company conducted an experiment and verification at a hydropower station in Chengdu. The total installed capacity of the reservoir is 3750kW, using three units. Event prediction is performed using the method described above. Assume the database output includes outflow (hereinafter referred to as I), tailwater level (hereinafter referred to as II), and dead water level (hereinafter referred to as III). Semantic information, such as flow rate, maximum output, total hydropower, energy consumption, minimum energy consumption, and error value, is obtained through an analyzer. Different classification label values ​​can be obtained from this semantic information. Then, a new semantic context is constructed using an n-grams model, such as obtaining reverse flow, forward flow, maximum processing time, and minimum processing time from the flow rate, maximum output, total hydropower, and energy consumption values. Event questions reflecting the hydropower station's operational data are obtained through this semantic approach. Finally, clustering and comprehensive evaluation models are used to analyze and calculate the obtained speech information in different forms, outputting data information such as the hydropower station's operating power.

[0118] In a further embodiment, assuming power generation is conducted under the influence of three factors, such as outflow, tailrace level, and dead water level, the total cost, equipment cost, and loss cost of the hydropower station are observed, in yuan. The output power of the hydropower station under these influencing factors is analyzed at different times. Then, based on the output power, the minimum cost within different time periods is calculated. Time period (unit: hours) information is shown in Table 1.

[0119] Table 1 Output power of hydropower stations at different time periods

[0120]

[0121] Based on the above description, after the semantic label of minimum power cost is identified by the above method, it is then calculated using the following formula, where f1-f7 represent the power output in different time periods. Specifically, minf1 = 200234; minf2 = 220214; minf3 = 210124; minf4 = 221100; minf5 = 202110; minf6 = 201410; minf7 = 211010. Therefore, according to the formula:

[0122] Then, by using the above method, multiple experiments were conducted to calculate the average value in order to improve the data accuracy, and the error data shown in Table 2 was output.

[0123] Table 2 Error Data Analysis

[0124]

[0125] After semantic recognition, the average error of the hydropower station's operating power data can be quickly identified, enabling rapid further analysis and calculation of the error data. Through the above calculations, the error rate is low when applying the mathematical model described above. The method of this application provides a technical basis for analyzing the economic operation of hydropower stations.

[0126] In a further embodiment, a decision tree similarity matrix is ​​established. Taking outflow volume, tailwater level, and dead water level as examples, the tree-structured decision tree is transformed into a binary tree structure, resulting in a set of {"outflow volume <= tailwater level", "tailwater level > dead water level", "dead water level <= outflow volume", etc.}.

[0127] The similarity Simn calculation formula is shown in Equation 1.

[0128]

[0129] Formula for calculating element similarity (Simn):

[0130]

[0131] Assuming data is extracted over 10 seconds, let t = 10, k1 = 10, k2 = 8, then substituting into formula (4), we get:

[0132]

[0133] Using the method described above, and taking different values ​​for N, calculate sim1-sim n For ease of calculation, let's assume that the different values ​​are 1.234, 1.302, 1.071, 1.321, 1.021, 1.001, 0.921, 0.673, 0.542, -1.86, etc., then... The different values ​​within are:

[0134]

[0135] Arbitrary similarity projection through matrix Different predicted values ​​can be output based on the magnitude of this value. This is determined by the prediction threshold; when the value is greater than the set threshold, it is considered abnormal data, and when it is less than the set threshold, it is considered normal data.

[0136] In a further specific embodiment, when using the LA algorithm again, assuming that when applying 2n event data parameters, the male dataset lionA m =[1, 2, 3, ..., 100] and female lionA f = [1, 2, 3, ..., 100]. When four new event data parameters A...1~4 Two randomly selected intersection points are used to generate four new event data parameters, intersection point A, through p-random mutation. 5~8 After crossover and mutation, a total of 8 new event data parameters are generated. The objective function f(y) of the LA algorithm is then substituted into... Then we can have:

[0137]

[0138] In calculating the optimal solution, through continuous calculation, when f(y) b )<f(y),y b When ≠ y, the optimal solution is output. Within the specified iteration interval, assume M... max The value is 1.46, P j It is 45. If the value is 294, then the cloning event data parameter M obtained from the location center of the ion population is:

[0139]

[0140] By cloning event data parameter M, other parameter formulas are calculated. This invention is merely illustrative of this embodiment and is not limited to the above embodiment. The data information in the above embodiment can be set according to actual measurement. Different datasets will result in different parameter sets and ultimately different calculated parameters. Generally, the larger the dataset, the more accurate the calculation. The embodiments of this invention are merely data-driven descriptions of the above theory; this invention is not limited to the description of the above embodiments, nor are these embodiments the only embodiments.

[0141] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are merely illustrative. Those skilled in the art can omit, substitute, and modify the details of the above methods and systems in various ways without departing from the principles and essence of the present invention. For example, combining the above method steps to perform substantially the same function and achieve substantially the same result according to substantially the same method falls within the scope of the present invention. Therefore, the scope of the present invention is defined only by the appended claims.

Claims

1. A method for event extraction and prediction based on temporal events and semantic context, characterized in that: Includes the following steps: (S1) Real-time acquisition of event data information, through image acquisition, video acquisition, and semantic background acquisition; The real-time acquisition of event data in step (S1) is based on an embedded multi-channel digital video acquisition device, including a video input interface, a video acquisition module, a core processor module, a video storage module, an external interface module, and a video transmission module. The output of the video input interface is connected to the input of the video acquisition module, the output of the video acquisition module is connected to the input of the core processor module, the output of the core processor module is connected to the input of the video storage module, the output of the core processor module is also connected to the input of the external interface module, and the video storage module is also connected to the video transmission module. The main video control chip uses either a TMS320DM8168 chip or a TVP5158 chip, and the main control chip includes an ARM module, a video processing module, an OCR recognition module, and a DSP module; the TVP5158 chip includes an FPGA module. (S2) Store the acquired video stream to obtain video stream data information; at the same time, generate a time-series event stream from the acquired video stream data information and record the timestamp; at the same time, convert the acquired data information into short text samples and record the short text sample labels; The method for generating a time-series event stream from the acquired video stream data information and recording timestamps is as follows: set the mVideo data acquisition interface, set the data stream input time and transmission time, construct the time-series event identification domain, and record the data information falling within the time-series event identification domain as a timestamp; The method for recording short text sample labels is a text similarity evaluation method, which adopts a text similarity method based on the n-grams model; (S3) Extract data features: Video stream data information is extracted using an image recognition model, and short text samples are analyzed using a constructed classifier; the data analysis method in step (S3) is the random forest algorithm. (S4) Construct a predictive event model and a semantic context model to achieve data prediction; In step (S4), the prediction model for the predicted event is constructed using the NILA-GCN model, which includes an input layer, a convolutional layer, a fusion layer, and a loss function module. The output of the input layer is connected to the input of the convolutional layer, the output of the convolutional layer is connected to the input of the fusion layer, and the output of the fusion layer is connected to the input of the loss function module. The semantic background model is constructed by a database, an analyzer, an n-grams model, a clustering model, and a comprehensive evaluation model. The output of the database is connected to the input of the analyzer, the output of the analyzer is connected to the input of the n-grams model, the output of the n-grams model is connected to the input of the clustering model, and the output of the clustering model is connected to the input of the comprehensive evaluation model. The prediction method of the semantic background model is as follows: (1) The system server collects semantic parameter topic data by outputting data information through the database; the system server submits the user's query terms to the search engine and allows the user to select the preferred semantics in the results of the returned page and form a new semantic set for the user; (2) Semantic information is obtained through the analyzer, and the system server establishes a new semantic model for the user; the system server establishes a concept map reflecting the user's preferences through the new semantic set; the system server first constructs a concept lattice before establishing the new semantic model for the user; (3) A new semantic background is constructed through the n-grams model, and the system server establishes a conceptual semantic background graph; the system server converts the concept lattice into a conceptual semantic background graph that can intuitively represent the semantic relationships between semantics; (4) The semantic information obtained is analyzed and calculated in different forms through clustering model and comprehensive evaluation model. The system server updates the concept semantic background map to update the new semantic data of users. The system server increases or decreases the concept semantic background map and finally outputs semantic information with high purity. (S5) Output the prediction results.

2. The event extraction and prediction method based on temporal events and semantic context according to claim 1, characterized in that: In step (S4), the prediction method of the prediction event prediction model is as follows: (S41) Population initialization, input 2 n The acquired video data parameters were evenly distributed to two groups as candidate populations, namely male lionA. m =[ y 1. y 2. y 3、•••、 y i ] and female lionA f =[ y 1. y 2. y 3、•••、 y i ]; (S42) Crossover and mutation; The acquisition of video data parameters generates new individuals through crossover and mutation. A crossover algorithm based on dual probabilities is used to realize the probabilities of two different data types. Male lionA m and female lionA f New video data parameter A was generated through mating. cub =[ y 1. y 2. y 3、•••、 y i ]; (S43) Territorial defense: In the process of iteratively generating superior individuals from the acquired video data parameters, male lionA m and female lionA f It will be attacked by lions with abnormal external information; at this time, male lionA m The goal is to protect and safeguard the high-quality video data parameters acquired, and to delineate the area surrounding the population as territory. (S44) yields the optimal solution, male lionA m and female lionA f Inferior solutions will be replaced with optimal solutions. Crossover will not end until the iteration termination condition is met. Once the iteration termination condition is met, crossover ends and the optimal solution is output.

3. The event extraction and prediction method based on temporal events and semantic context according to claim 1, characterized in that: The method for recording the timestamp in step (S2) is as follows: Configure the mVideo data acquisition interface, set the data stream input time and transmission time, construct a time-series event identification domain, and record data information falling within the time-series event identification domain as a timestamp. The time-series event identification domain construction process involves comparing historical time-series event data with the model, then calculating the difference between the historical data and the model data. If over 96% of the perceived video data information points fall within the area surrounding the model, then the formula is activated. (1); When formula (1) is a polynomial equation that approximates the real time sequence event, the principle of formula (1) is the principle of linear algebra, which connects different data information through linear connection to represent the data information of the perceived video data information points.

4. The event extraction and prediction method based on temporal events and semantic context according to claim 1, characterized in that: The text similarity method based on the n-grams model is as follows: First, all semantic texts related to the event are retrieved from the text database. Then, appropriate preprocessing is performed to obtain an n-grams semantic model. The numerical expression or frequency matrix of the text is found for visualization, clustering, and numerical estimation. After creating the frequency matrix, the matrix is ​​provided to the SOM. The dataset is clustered and visualized in the SOM to detect the visual similarity of the texts. At the same time, four algorithms, namely the cosine similarity method, the dice similarity method, the extended Jaccard similarity method, and the overlap coefficient similarity method, are used to calculate the similarity measure or numerical estimate. In the process of creating the frequency matrix, the final calculation result depends on the creation of n-gram words, including the selected filters. n-gram words include sentences, paragraphs, keywords, and semantics.

Citation Information

Patent Citations

  • Video content recognition method and device, storage medium and electronic equipment

    CN113569610A