Long video content information acquisition method based on OCR (Optical Character Recognition) and voice recognition technology
By combining the horned lizard optimization algorithm and the belief propagation algorithm, the problems of unstable recognition accuracy and inconsistent modal semantics in long videos are solved, efficient multimodal information collection and fusion are achieved, and the recognition accuracy and system efficiency are improved.
Patent Information
- Application Number
- CN202510755057.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-12
AI Technical Summary
Existing multimodal information acquisition technologies have problems in long video processing scenarios, such as unstable recognition accuracy, inconsistent modal semantics, and low system operation efficiency. They also lack the capabilities of adaptive parameter optimization and multimodal semantic fusion.
The horned lizard optimization algorithm is combined with the belief propagation algorithm to construct a multi-objective fitness function to adaptively optimize the OCR and ASR recognition parameters, and semantic consistency fusion is performed through the factor graph structure to achieve efficient extraction and fusion of image text and speech text.
It significantly improves the recognition accuracy and semantic consistency of image text and voice text in long videos, reduces processing latency and resource consumption, improves the collaborative processing capability of multimodal information, and enhances the degree of automation of video understanding.
Smart Images

Figure CN120635776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimodal intelligent information processing technology, and in particular to a method for collecting long video content information based on OCR and speech recognition technologies. Background Art
[0002] With the rapid development of artificial intelligence (AI), technologies for automatic recognition and structured processing of video content have been widely applied in a variety of fields, including public opinion monitoring, educational resource compilation, judicial record analysis, and media content management. Long videos, as important information carriers, typically contain a large amount of multimodal content, including images, audio, and subtitles. The efficient and accurate extraction and integration of valuable information from these content has become a hot topic of current research. Existing mainstream methods typically use OCR (Optical Character Recognition) technology to perform text recognition on video image frames, combined with ASR (Automatic Speech Recognition) technology to transcribe audio streams, thereby achieving a preliminary restoration of the semantic content in the video.
[0003] However, existing multimodal information acquisition technologies still have many limitations when applied to long video processing scenarios. First, the OCR and ASR recognition parameter settings mostly rely on static configuration or manual experience, and lack adaptive adjustment to the specific video content structure and quality changes, resulting in a significant decrease in recognition accuracy in complex backgrounds, low resolution or noisy speech segments. Secondly, there are semantic offsets and time synchronization errors in the recognition results of different modalities. Traditional simple splicing or scoring fusion mechanisms are difficult to ensure information consistency, which reduces the accuracy of the final semantic analysis. Thirdly, in actual scenarios, the video content is long, the modal information is dense and changes dramatically. How to balance system operation efficiency and resource consumption while ensuring recognition quality is still a bottleneck in technological development.
[0004] Furthermore, some studies have attempted to introduce heuristic algorithms for parameter optimization, but most have been applied only to single-modal tasks and lack systematic modeling for the coordinated optimization of OCR and ASR parameters. Furthermore, for multimodal semantic fusion, most have employed fixed rules or shallow attention mechanisms, yet have yet to implement structural causal reasoning and joint confidence judgment. Consequently, existing technologies have yet to develop a comprehensive video information acquisition framework that combines adaptive parameter optimization, multimodal semantic consistency reasoning, and structured output.
[0005] Therefore, how to provide a long video content information collection method based on OCR and speech recognition technology is a problem that technicians in this field urgently need to solve. Summary of the Invention
[0006] One objective of the present invention is to propose a method for collecting content information from long videos based on optical character recognition (OCR) and speech recognition technologies. This method integrates the Horned Lizard optimization algorithm with the belief propagation algorithm to adaptively optimize multimodal recognition parameters and integrate semantic consistency for image frame sequences and audio streams in long videos. Leveraging technologies such as multi-objective intelligent optimization, factor graph reasoning, and deep text embedding, the method describes in detail how to efficiently extract image and speech text information from complex video scenes, constructing a unified semantic content information set. This method boasts high recognition accuracy, strong modal fusion capabilities, and robust processing.
[0007] According to an embodiment of the present invention, a method for collecting long video content information based on OCR and speech recognition technology includes the following steps:
[0008] S1. Preprocess the input long video data, segment it according to a fixed time window, and extract the video image frame sequence and corresponding audio stream of each segment;
[0009] S2. Input the video image frame sequence into the OCR recognition module for text recognition processing, input the synchronously extracted audio stream into the ASR recognition module for speech recognition processing, set the initial parameter group, and obtain preliminary recognition results;
[0010] S3. Based on the accuracy, processing delay, resource consumption, and inter-modal confidence consistency of the preliminary recognition results, a multi-objective fitness function is constructed to initialize the individual population of the horned lizard optimization algorithm, generate and update individual solutions containing multiple OCR and ASR parameter combinations, and finally select the parameter group with the best fitness;
[0011] S4, applying the optimal parameter group to the OCR and ASR recognition modules respectively, performing recognition processing on all video segments, and obtaining optimized image text information and voice text information;
[0012] S5. Based on the obtained recognition results, a fusion factor graph structure is constructed, and the belief propagation algorithm is used to perform edge message passing and joint reasoning to obtain a set of multimodal semantic blocks;
[0013] S6. Extract, merge, and structure the multimodal semantic block set to generate a unified multimodal content information set for video content retrieval, index construction, or semantic analysis.
[0014] Optionally, segmentation is performed according to a fixed time window, and the video image frame sequence and the corresponding audio stream of each segment are extracted, wherein the time window size is 10 seconds, and the image frame sequence and the corresponding audio stream are extracted in each segment.
[0015] Optionally, the S2 specifically includes:
[0016] S21, inputting the extracted image frame sequence into the OCR recognition module, performing size standardization, grayscale enhancement and noise suppression on the image frames, and detecting text candidate regions through a neural network model;
[0017] S22, setting the frame extraction interval in the initial parameter group to Δf, the image clarity threshold to τ1, and the image text detection confidence threshold to τ2, where Δf means extracting a frame every Δf frames, τ1 represents the minimum acceptable clarity score of the image frame, and τ2 represents the minimum confidence required for the candidate region to be identified as a text region;
[0018] S23, will satisfy τ1≥τ 1_min And τ2≥τ 2_min The image frame with the required conditions is input to the character recognition submodule of the OCR recognition module to complete the text decoding and output the image text information set T img ={t1,t2,…,t n}, where t i Represents the text content recognized in the i-th frame, i∈[1,n], where n is the total number of valid image frames in the current video segment after parameter screening, τ 1_min Indicates the minimum acceptable threshold of image clarity, τ 2_min Indicates the minimum confidence threshold for a candidate text region to be considered a valid text region;
[0019] S24, inputting the extracted audio stream into the ASR recognition module, performing pre-emphasis, endpoint detection and spectrum conversion on the audio, and extracting Mel-frequency cepstral coefficients as audio feature vectors;
[0020] S25. Set the audio slice length in the initial parameter group to L1, the frame shift length to L2, and the language model fusion ratio to α, where L1∈[500,5000] represents the length of each speech slice, L2∈[10,50] represents the overlap interval between slices, and α represents the decoding fusion ratio of the language model in the ASR recognition module;
[0021] S26: Input the audio feature vector into the acoustic model and language model of the ASR recognition module for joint decoding, and output the speech text information set T speech ={s1,s2,…,s m}, where s j Represents the text content recognized in the jth speech slice, j∈[1,m], where m is the number of valid speech slices segmented in the current video segment;
[0022] S27, the image text information set T img and the voice text information set T speech As a preliminary identification result.
[0023] Optionally, the S3 specifically includes:
[0024] S31. Based on the obtained preliminary recognition results, construct a multi-objective evaluation index, which includes the recognition accuracy of image text information and voice text information, processing delay, resource consumption, inter-modal confidence consistency score, and edge stability parameter of the confidence propagation output result, to form an evaluation feature vector set Z (k) ;
[0025] S32, setting the population size of the horned lizard optimization algorithm to N, and initializing an initial population containing multiple individual parameter combinations, each individual parameter combination represents a set of candidate recognition parameters of the OCR recognition module and the ASR recognition module, including the image frame frequency difference Δf (k) , image clarity threshold Text confidence threshold Minimum size threshold for text candidate regions Minimum length of voice slice window The correlation coefficient threshold α with the audio semantic segment (k) ;
[0026] S33. Set up a direction memory mechanism for each individual, maintain the direction angle change and inter-modal confidence consistency score record of the individual in the past h iterations in each round of iteration, and construct a direction memory vector
[0027] S34. During the individual parameter update process, the individual direction angle is updated using an angle perturbation strategy with a gradient term, in combination with the consistency change trend recorded in the direction memory vector:
[0028]
[0029] in, represents the mean value of modal consistency in the past h times, γ is the direction memory adjustment factor, is the direction angle of the k-th horned lizard individual at the t-th iteration, is the direction angle of the kth horned lizard individual at the t-1th iteration, β is the angle perturbation factor, is a standard normally distributed random number, is the gradient of modal consistency change in the memory direction;
[0030] S35. Calculate the local position distribution entropy H of the current horned lizard individual population (t) , evaluate the population convergence state and diversity level, the local position distribution entropy is used to dynamically adjust the individual angle disturbance factor β (t) and jump step factor J (t) ;
[0031] S36. According to the jump update strategy, in the current iteration round, a jump update is performed based on the individual's current parameter vector, the optimal individual position and direction angle:
[0032]
[0033] in, is the updated parameter combination vector of the kth individual in the t+1th round of iteration, is the current parameter combination vector of the kth horned lizard individual in the tth iteration, J (t) is the jump step coefficient in the tth iteration, cos is the cosine function, is the individual parameter combination with the highest fitness score in the tth iteration;
[0034] S37. In each iteration, each updated parameter combination is applied to the OCR recognition module and the ASR recognition module respectively, and the extracted image frame sequence and audio stream are recognized and processed to obtain the image text information set respectively. Voice and text information collection
[0035] S38. Based on the recognition results of the image and speech, calculate the recognition accuracy, processing delay, resource consumption, modal confidence consistency and confidence propagation edge stability of the individual parameter combination in the current round, and combine the above eigenvalues into a eigenvector Z (k) , input into the joint evaluation network to obtain the individual fitness score F (k) ;
[0036] S39, perform iterative fitness evaluation on each individual, compare the fitness score of this round with the score of the previous round, if F (k,t+1) >F (k,t) , then keep the new parameter combination and update the direction memory vector Otherwise, keep the original parameters unchanged and update the individual direction disturbance state, where F (k,t+1) In the horned lizard optimization algorithm, the fitness score of the kth individual in the t+1th iteration, F (k,t) represents the fitness score of the same individual in the tth iteration;
[0037] S310, repeat steps S33 to S39 until the set maximum number of iterations T is reached max Or the global optimal fitness change satisfies the convergence judgment threshold ε, and finally outputs the parameter combination X with the optimal fitness score * .
[0038] Optionally, the S4 specifically includes:
[0039] S41, the optimal parameter combination vector output by the horned lizard optimization algorithm They are respectively configured into the OCR recognition module and the ASR recognition module as the final recognition parameters. The OCR recognition module is applied to the video image frame sequence, and the ASR recognition module is applied to the audio stream data, where Δf * is the optimized image frame frequency difference, is the optimized image clarity threshold, is the optimized text confidence threshold is the minimum size threshold of the optimized text candidate area, is the minimum length of the optimized speech slice window, α * is the optimized audio semantic segment correlation coefficient threshold;
[0040] S42, for all the segmented video contents obtained by preprocessing, call the OCR recognition module and the ASR recognition module segment by segment according to the segment number i, and use the optimal parameter combination vector X * Perform identification processing;
[0041] S43, in each segment i, the OCR recognition module uses the parameter Δf * , Process the image frame sequence, extract the text region candidate box, confidence score and image character results, and obtain the image text information set;
[0042] S44, in each segment i, the ASR recognition module uses parameters α * Slice the audio stream and extract voiceprint features, and perform speech-to-text conversion to obtain a speech-to-text information set;
[0043] S45, the image text information set obtained for each segment i Voice and text information collection Perform preliminary alignment based on timestamps. The alignment strategy matches the appearance frame position of the text segment with the time window overlap of the speech slice to form a segmented fusion text set.
[0044] S46, merging the fused text sets of all video segments i to obtain the overall optimized recognized image text information and voice text information;
[0045] S47. Send the merged image text information and speech text information to the belief propagation processing module respectively to perform semantic reasoning, fusion and final information structure generation.
[0046] Optionally, the S5 specifically includes:
[0047] S51, based on the acquired image text information set T imgand the voice text information set T speech , construct the modality fusion factor graph structure G = (V, F, E), where V represents all modality semantic variable nodes to be jointly reasoned, including the image segment text variable node v i With the speech segment text variable node v j , F represents the set of semantic relationship factors between modalities, and E represents the set of edges connecting semantic variable nodes and relationship factors;
[0048] S52, constructing a set of modal factor functions Φ={φ ij}, used to model the semantic relevance scoring function between image text segments and speech text segments, defined as:
[0049]
[0050] Among them, Emb(v i ) represents the text fragment v i FastText embedding vector, σ is the smoothing factor that adjusts the influence range of the embedding space distance, Emb(v j ) represents the text fragment v j FastText embedding vector, exp is the exponential function;
[0051] S53, for each variable node v k Initialize the prior marginal probability distribution Normalized generation is performed based on the confidence score in the OCR recognition module or the ASR recognition module;
[0052] S54, for the constructed factor graph G, the variational belief propagation algorithm is used to iteratively calculate the edge probability, and for each connecting edge e ij , execute the message passing rules:
[0053]
[0054] in, represents the message transmitted from node i to node j in round t+1, N(i) represents the set of adjacent nodes of node i, represents the message transmitted from node i to node j in round t;
[0055] S55. After each round of iteration, the edge probabilities of all variable nodes are updated using a normalization correction strategy:
[0056]
[0057] Among them, Z j is the normalization factor, Indicates that in the t+1th iteration, the variable node v in the factor graph jThe marginal probability distribution function of
[0058] S56, repeat steps S54 to S55 until the edge probabilities of all variable nodes converge or the maximum number of iterations T is reached. BP , at this time the final marginal probability distribution As the joint semantic confidence output of the text fragments, according to the joint semantic confidence output, the text fragments in the image text and the speech text with confidence higher than the set threshold θ are marked as semantic blocks to form a multimodal semantic block set.
[0059] The beneficial effects of the present invention are:
[0060] The present invention introduces the horned lizard optimization algorithm and the belief propagation algorithm to achieve adaptive adjustment of recognition parameters and deep fusion of multimodal information based on OCR and ASR recognition technologies, effectively overcoming the shortcomings of existing technologies in terms of fixed parameter configuration, unstable recognition accuracy and inconsistent modal semantics. The horned lizard optimization algorithm is used to construct a multi-objective fitness function, and multiple indicators such as recognition accuracy, processing delay, resource consumption and confidence consistency between modalities are jointly optimized, which significantly improves the recognition effect and processing efficiency in complex video clips. The factor graph structure constructed by the belief propagation algorithm can model and infer the potential semantic relationship between image text and speech text, realize the information collaborative judgment and consistency enhancement between modalities, and avoid the semantic offset and redundant interference problems that are prone to occur in traditional weighted fusion methods.
[0061] In addition, the present invention fully considers the segmentation characteristics and modal concurrency features of long video processing in the overall framework design, and adopts a combination of fixed time window segmentation and nested recognition strategy, which makes the method have good scalability and engineering adaptability. The final generated multimodal semantic block set can directly serve downstream tasks such as video content retrieval, index construction and semantic analysis after structured processing, enhancing the degree of automation of video understanding and data utilization efficiency. In summary, the present invention significantly improves the level of intelligent processing and the ability to collaboratively process multimodal information while maintaining recognition accuracy, and has strong practical application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0063] Figure 1 This is a flow chart of the method for collecting long video content information based on OCR and speech recognition technology proposed by the present invention;
[0064] Figure 2A processing flow chart for optimizing the multi-objective fitness function and outputting the optimal recognition parameter combination using the horned lizard optimization algorithm of the long video content information collection method based on OCR and speech recognition technology proposed in the present invention. DETAILED DESCRIPTION
[0065] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0066] refer to Figure 1 and Figure 2 , a method for collecting long video content information based on OCR and speech recognition technology, comprising the following steps:
[0067] S1. Preprocess the input long video data, segment it according to a fixed time window, and extract the video image frame sequence and corresponding audio stream of each segment;
[0068] S2. Input the video image frame sequence into the OCR recognition module for text recognition processing, input the synchronously extracted audio stream into the ASR recognition module for speech recognition processing, set the initial parameter group, and obtain preliminary recognition results;
[0069] S3. Based on the accuracy, processing delay, resource consumption, and inter-modal confidence consistency of the preliminary recognition results, a multi-objective fitness function is constructed to initialize the individual population of the horned lizard optimization algorithm, generate and update individual solutions containing multiple OCR and ASR parameter combinations, and finally select the parameter group with the best fitness;
[0070] S4, applying the optimal parameter group to the OCR and ASR recognition modules respectively, performing recognition processing on all video segments, and obtaining optimized image text information and voice text information;
[0071] S5. Based on the obtained recognition results, a fusion factor graph structure is constructed, and the belief propagation algorithm is used to perform edge message passing and joint reasoning to obtain a set of multimodal semantic blocks;
[0072] S6. Extract, merge, and structure the multimodal semantic block set to generate a unified multimodal content information set for video content retrieval, index construction, or semantic analysis.
[0073] In this embodiment, the video image frame sequence and the corresponding audio stream of each segment are extracted by segmenting according to a fixed time window, wherein the time window size is 10 seconds, and the image frame sequence and the corresponding audio stream are extracted in each segment.
[0074] In this embodiment, S2 specifically includes:
[0075] S21, inputting the extracted image frame sequence into the OCR recognition module, performing size standardization, grayscale enhancement and noise suppression on the image frames, and detecting text candidate regions through a neural network model;
[0076] S22, setting the frame extraction interval in the initial parameter group to Δf, the image clarity threshold to τ1, and the image text detection confidence threshold to τ2, where Δf means extracting a frame every Δf frames, τ1 represents the minimum acceptable clarity score of the image frame, and τ2 represents the minimum confidence required for the candidate region to be identified as a text region;
[0077] S23, will satisfy τ1≥τ 1_min And τ2≥τ 2_min The image frame with the required conditions is input to the character recognition submodule of the OCR recognition module to complete the text decoding and output the image text information set T img ={t1,t2,…,t n}, where t i Represents the text content recognized in the i-th frame, i∈[1,n], where n is the total number of valid image frames in the current video segment after parameter screening, τ 1_min Indicates the minimum acceptable threshold of image clarity, τ 2_min Indicates the minimum confidence threshold for a candidate text region to be considered a valid text region;
[0078] S24, inputting the extracted audio stream into the ASR recognition module, performing pre-emphasis, endpoint detection and spectrum conversion on the audio, and extracting Mel-frequency cepstral coefficients as audio feature vectors;
[0079] S25. Set the audio slice length in the initial parameter group to L1, the frame shift length to L2, and the language model fusion ratio to α, where L1∈[500,5000] represents the length of each speech slice, L2∈[10,50] represents the overlap interval between slices, and α represents the decoding fusion ratio of the language model in the ASR recognition module;
[0080] S26: Input the audio feature vector into the acoustic model and language model of the ASR recognition module for joint decoding, and output the speech text information set T speech ={s1,s2,…,s m}, where s j Represents the text content recognized in the jth speech slice, j∈[1,m], where m is the number of valid speech slices segmented in the current video segment;
[0081] S27, the image text information set T img and the voice text information set T speech As a preliminary identification result.
[0082] In this embodiment, S3 specifically includes:
[0083] S31. Based on the obtained preliminary recognition results, construct a multi-objective evaluation index, which includes the recognition accuracy of image text information and voice text information, processing delay, resource consumption, inter-modal confidence consistency score, and edge stability parameter of the confidence propagation output result, to form an evaluation feature vector set Z (k) ;
[0084] S32, setting the population size of the horned lizard optimization algorithm to N, and initializing an initial population containing multiple individual parameter combinations, each individual parameter combination represents a set of candidate recognition parameters of the OCR recognition module and the ASR recognition module, including the image frame frequency difference Δf (k) , image clarity threshold Text confidence threshold Minimum size threshold for text candidate regions Minimum length of voice slice window The correlation coefficient threshold α with the audio semantic segment (k) ;
[0085] S33. Set up a direction memory mechanism for each individual, maintain the direction angle change and inter-modal confidence consistency score record of the individual in the past h iterations in each round of iteration, and construct a direction memory vector
[0086] S34. During the individual parameter update process, the individual direction angle is updated using an angle perturbation strategy with a gradient term, in combination with the consistency change trend recorded in the direction memory vector:
[0087]
[0088] in, represents the mean value of modal consistency in the past h times, γ is the direction memory adjustment factor, is the direction angle of the k-th horned lizard individual at the t-th iteration, is the direction angle of the kth horned lizard individual at the t-1th iteration, β is the angle perturbation factor, is a standard normally distributed random number, is the gradient of modal consistency change in the memory direction;
[0089] S35. Calculate the local position distribution entropy H of the current horned lizard individual population (t) , evaluate the population convergence state and diversity level, the local position distribution entropy is used to dynamically adjust the individual angle disturbance factor β (t) and jump step factor J (t) ;
[0090] S36. According to the jump update strategy, in the current iteration round, a jump update is performed based on the individual's current parameter vector, the optimal individual position and direction angle:
[0091]
[0092] in, is the updated parameter combination vector of the kth individual in the t+1th round of iteration, is the current parameter combination vector of the kth horned lizard individual in the tth iteration, J (t) is the jump step coefficient in the tth iteration, cos is the cosine function, is the individual parameter combination with the highest fitness score in the tth iteration;
[0093] S37. In each iteration, each updated parameter combination is applied to the OCR recognition module and the ASR recognition module respectively, and the extracted image frame sequence and audio stream are recognized and processed to obtain the image text information set respectively. Voice and text information collection
[0094] S38. Based on the recognition results of the image and speech, calculate the recognition accuracy, processing delay, resource consumption, modal confidence consistency and confidence propagation edge stability of the individual parameter combination in the current round, and combine the above eigenvalues into a eigenvector Z (k) , input into the joint evaluation network to obtain the individual fitness score F (k) ;
[0095] S39, perform iterative fitness evaluation on each individual, compare the fitness score of this round with the score of the previous round, if F (k,t+1) >F (k,t) , then keep the new parameter combination and update the direction memory vector Otherwise, keep the original parameters unchanged and update the individual direction disturbance state, where F (k,t+1) In the horned lizard optimization algorithm, the fitness score of the kth individual in the t+1th iteration, F (k,t) represents the fitness score of the same individual in the tth iteration;
[0096] S310, repeat steps S33 to S39 until the set maximum number of iterations T is reached max Or the global optimal fitness change satisfies the convergence judgment threshold ε, and finally outputs the parameter combination X with the optimal fitness score * .
[0097] In this embodiment, the S4 specifically includes:
[0098] S41, the optimal parameter combination vector output by the horned lizard optimization algorithm They are respectively configured into the OCR recognition module and the ASR recognition module as the final recognition parameters. The OCR recognition module is applied to the video image frame sequence, and the ASR recognition module is applied to the audio stream data, where Δf * is the optimized image frame frequency difference, is the optimized image clarity threshold, is the optimized text confidence threshold is the minimum size threshold of the optimized text candidate area, is the minimum length of the optimized speech slice window, α * is the optimized audio semantic segment correlation coefficient threshold;
[0099] S42, for all the segmented video contents obtained by preprocessing, call the OCR recognition module and the ASR recognition module segment by segment according to the segment number i, and use the optimal parameter combination vector X * Perform identification processing;
[0100] S43, in each segment i, the OCR recognition module uses the parameter Δf * , Process the image frame sequence, extract the text region candidate box, confidence score and image character results, and obtain the image text information set;
[0101] S44, in each segment i, the ASR recognition module uses parameters α * Slice the audio stream and extract voiceprint features, and perform speech-to-text conversion to obtain a speech-to-text information set;
[0102] S45, the image text information set obtained for each segment i Voice and text information collection Perform preliminary alignment based on timestamps. The alignment strategy matches the appearance frame position of the text segment with the time window overlap of the speech slice to form a segmented fusion text set.
[0103] S46, merging the fused text sets of all video segments i to obtain the overall optimized recognized image text information and voice text information;
[0104] S47. Send the merged image text information and speech text information to the belief propagation processing module respectively to perform semantic reasoning, fusion and final information structure generation.
[0105] In this embodiment, the S5 specifically includes:
[0106] S51, based on the acquired image text information set Timg and the voice text information set T speech , construct the modality fusion factor graph structure G = (V, F, E), where V represents all modality semantic variable nodes to be jointly reasoned, including the image segment text variable node v i With the speech segment text variable node v j , F represents the set of semantic relationship factors between modalities, and E represents the set of edges connecting semantic variable nodes and relationship factors;
[0107] S52, constructing a set of modal factor functions Φ={φ ij}, used to model the semantic relevance scoring function between image text segments and speech text segments, defined as:
[0108]
[0109] Among them, Emb(v i ) represents the text fragment v i FastText embedding vector, σ is the smoothing factor that adjusts the influence range of the embedding space distance, Emb(v j ) represents the text fragment v j FastText embedding vector, exp is the exponential function;
[0110] S53, for each variable node v k Initialize the prior marginal probability distribution Normalized generation is performed based on the confidence score in the OCR recognition module or the ASR recognition module;
[0111] S54, for the constructed factor graph G, the variational belief propagation algorithm is used to iteratively calculate the edge probability, and for each connecting edge e ij , execute the message passing rules:
[0112]
[0113] in, represents the message transmitted from node i to node j in round t+1, N(i) represents the set of adjacent nodes of node i, represents the message transmitted from node i to node j in round t;
[0114] S55. After each round of iteration, the edge probabilities of all variable nodes are updated using a normalization correction strategy:
[0115]
[0116] Among them, Z j is the normalization factor, Indicates that in the t+1th iteration, the variable node v in the factor graphj The marginal probability distribution function of
[0117] S56, repeat steps S54 to S55 until the edge probabilities of all variable nodes converge or the maximum number of iterations T is reached. BP , at this time the final marginal probability distribution As the joint semantic confidence output of the text fragments, according to the joint semantic confidence output, the text fragments in the image text and the speech text with confidence higher than the set threshold θ are marked as semantic blocks to form a multimodal semantic block set.
[0118] Example 1:
[0119] In order to verify the feasibility of the present invention in implementation, the present invention was applied to the online teaching project of the "Basics of Artificial Intelligence" course in a certain university. The academic affairs department hopes to automatically extract keywords, explanation paragraphs and core semantics from the long video courses recorded by teachers to build a course knowledge index and auxiliary question-answering system. However, the content of teaching videos is usually long (45 to 75 minutes) and contains three main information modes: blackboard writing, PPT explanation and teacher oral narration. The traditional OCR and ASR recognition systems using fixed parameters have poor processing effects: OCR is easily affected by factors such as blurred blackboard writing and light interference, and ASR has a high recognition error rate in a noisy environment. In addition, the output results of the two are often out of sync in time, semantics are repeated or lost, which seriously restricts the subsequent semantic structured processing and retrieval services.
[0120] In order to solve these problems, this embodiment adopts the method proposed in the present invention to perform structured collection and analysis on the above-mentioned teaching videos. The video is first divided into short clips with a window of 10 seconds, and the image frame sequence and audio stream are extracted separately. The system calls the horned lizard optimization algorithm to construct a multi-objective fitness function, and optimizes the recognition parameters by taking recognition accuracy, resource consumption, processing delay and semantic confidence consistency between OCR / ASR modalities as joint goals. During the optimization process, the initial population automatically generated more than 300 sets of parameter combinations, and finally converged to a set of optimal solutions, successfully improving the quality of text and speech extraction under conditions of uneven lighting and changing speech speed.
[0121] The optimized recognition results are fed into the belief propagation module, which constructs a factor graph structure with image and speech text as nodes, dynamically builds edge weights using text embedding semantic similarity, and performs message passing and joint reasoning. The system accurately merges multimodal segments, such as "Model training process (image)" and "Let's now look at the training process (speech)," that differ in expression but share the same semantics, into unified semantic blocks. Keywords are further extracted from these semantic blocks, and a knowledge fragment index is constructed. This information is then synchronized with the course management system for student retrieval.
[0122] During the system's online testing, 20 instructional videos from three courses, "Foundations of Artificial Intelligence," "Pattern Recognition," and "Introduction to Computer Vision," were processed. The experiment was conducted at a university from April 1 to April 15, 2025. Participants included three teaching assistants and 12 undergraduate student volunteers. When students used the system to retrieve after-class questions and locate content, the average time spent was reduced by approximately 60%, and the error rate was reduced by more than half. Teaching assistants reported a significant improvement in content extraction accuracy and semantic coherence, resulting in an overall system satisfaction rating of over 92%.
[0123] Table 1 Comparison of recognition and fusion performance between traditional methods and the method of the present invention
[0124]
[0125]
[0126] It can be clearly seen from Table 1 above that the present invention has achieved significant improvements in many key performance indicators. First, in terms of OCR recognition accuracy, the average value of the traditional method is 82.7%, while the method of the present invention introduces the horned lizard optimization algorithm to adaptively adjust the recognition parameters, thereby increasing the accuracy to 93.4%, an increase of 10.7%, and effectively solving the recognition difficulties such as blurred blackboard writing and image jitter. In terms of ASR recognition accuracy, it has also increased from the original 84.1% to 92.2%, thanks to the good adaptability of the parameter self-adjustment mechanism to changes in speech speed and voice environment, making the voice content restoration clearer and more accurate.
[0127] Regarding the consistency matching rate of multimodal information fusion, this paper constructs a factor graph and introduces a belief propagation algorithm for semantic reasoning, significantly improving the semantic alignment rate between OCR and ASR text from 68.9% with traditional methods to 89.5%, a 20.6% improvement. This effectively addresses the issues of inter-modal expression offset and temporal inconsistency. This improvement lays a foundation for higher-quality subsequent content extraction and structured analysis.
[0128] From a user experience perspective, the average content location response time has been significantly reduced from 58.6 seconds using traditional methods to 21.3 seconds, a reduction of approximately 63.6%, significantly improving user efficiency in information retrieval and knowledge retrieval scenarios. Subjective satisfaction scores have also increased from 3.6 to 4.7, indicating a greater appreciation for the accuracy and structure of the system's output.
[0129] Furthermore, in terms of the number of erroneous segments, traditional methods produce an average of 14.2 content recognition errors per hour of video, while the proposed method effectively reduces this to 5.6, a reduction of over 60%. This result fully demonstrates the comprehensive improvement of recognition robustness and information reliability achieved by the proposed method in practical applications.
[0130] In summary, the tabular data fully verifies the comprehensive advantages of the present invention in terms of accuracy, response efficiency, multimodal fusion consistency, and user experience, and it has high practicality and engineering promotion value.
[0131] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for collecting long video content information based on OCR and speech recognition technology, characterized in that: The steps include: S1. Preprocess the input long video data, segment it according to a fixed time window, and extract the video image frame sequence and corresponding audio stream of each segment; S2. Input the video image frame sequence into the OCR recognition module for text recognition processing, input the synchronously extracted audio stream into the ASR recognition module for speech recognition processing, set the initial parameter group, and obtain preliminary recognition results; S3. Based on the accuracy, processing delay, resource consumption, and inter-modal confidence consistency of the preliminary recognition results, a multi-objective fitness function is constructed to initialize the individual population of the horned lizard optimization algorithm, generate and update individual solutions containing multiple OCR and ASR parameter combinations, and finally select the parameter group with the best fitness; S4, applying the optimal parameter group to the OCR and ASR recognition modules respectively, performing recognition processing on all video segments, and obtaining optimized image text information and voice text information; S5. Based on the obtained recognition results, a fusion factor graph structure is constructed, and the belief propagation algorithm is used to perform edge message passing and joint reasoning to obtain a set of multimodal semantic blocks; S6. Extract, merge, and structure the multimodal semantic block set to generate a unified multimodal content information set for video content retrieval, index construction, or semantic analysis.
2. The method for collecting long video content information based on OCR and speech recognition technology according to claim 1, characterized in that: The video is segmented according to a fixed time window, and a video image frame sequence and a corresponding audio stream are extracted from each segment, wherein the time window size is 10 seconds, and an image frame sequence and a corresponding audio stream are extracted from each segment.
3. The method for collecting long video content information based on OCR and speech recognition technology according to claim 1, characterized in that: The S2 specifically includes: S21, inputting the extracted image frame sequence into the OCR recognition module, performing size standardization, grayscale enhancement and noise suppression on the image frames, and detecting text candidate regions through a neural network model; S22, setting the frame extraction interval in the initial parameter group to Δf, the image clarity threshold to τ1, and the image text detection confidence threshold to τ2, where Δf means extracting a frame every Δf frames, τ1 represents the minimum acceptable clarity score of the image frame, and τ2 represents the minimum confidence required for the candidate region to be identified as a text region; S23, will satisfy τ1≥τ 1_min And τ2≥τ 2_min The image frame with the required conditions is input to the character recognition submodule of the OCR recognition module to complete the text decoding and output the image text information set T img ={t1,t2,…,t n }, where t i Represents the text content recognized in the i-th frame, i∈[1,n], where n is the total number of valid image frames in the current video segment after parameter screening, τ 1_min Indicates the minimum acceptable threshold of image clarity, τ 2_min Indicates the minimum confidence threshold for a candidate text region to be considered a valid text region; S24, inputting the extracted audio stream into the ASR recognition module, performing pre-emphasis, endpoint detection and spectrum conversion on the audio, and extracting Mel-frequency cepstral coefficients as audio feature vectors; S25. Set the audio slice length in the initial parameter group to L1, the frame shift length to L2, and the language model fusion ratio to α, where L1∈[500,5000] represents the length of each speech slice, L2∈[10,50] represents the overlap interval between slices, and α represents the decoding fusion ratio of the language model in the ASR recognition module; S26: Input the audio feature vector into the acoustic model and language model of the ASR recognition module for joint decoding, and output the speech text information set T speech ={s1,s2,…,s m }, where s j Represents the text content recognized in the jth speech slice, j∈[1,m], where m is the number of valid speech slices segmented in the current video segment; S27, the image text information set T img and the voice text information set T speech As a preliminary identification result.
4. The method for collecting long video content information based on OCR and speech recognition technology according to claim 1, characterized in that: The S3 specifically includes: S31. Based on the obtained preliminary recognition results, construct a multi-objective evaluation index, which includes the recognition accuracy of image text information and voice text information, processing delay, resource consumption, inter-modal confidence consistency score, and edge stability parameter of the confidence propagation output result, to form an evaluation feature vector set Z (k) ; S32, setting the population size of the horned lizard optimization algorithm to N, and initializing an initial population containing multiple individual parameter combinations, each individual parameter combination represents a set of candidate recognition parameters of the OCR recognition module and the ASR recognition module, including the image frame frequency difference Δf (k) , image clarity threshold Text confidence threshold Minimum size threshold for text candidate regions Minimum length of voice slice window The correlation coefficient threshold α with the audio semantic segment (k) ; S33. Set up a direction memory mechanism for each individual, maintain the direction angle change and inter-modal confidence consistency score record of the individual in the past h iterations in each round of iteration, and construct a direction memory vector S34. During the individual parameter update process, the individual direction angle is updated using an angle perturbation strategy with a gradient term, in combination with the consistency change trend recorded in the direction memory vector: in, represents the mean value of modal consistency in the past h times, γ is the direction memory adjustment factor, is the direction angle of the k-th horned lizard individual at the t-th iteration, is the direction angle of the kth horned lizard individual at the t-1th iteration, β is the angle perturbation factor, is a standard normally distributed random number, is the gradient of modal consistency change in the memory direction; S35. Calculate the local position distribution entropy H of the current horned lizard individual population (t) , evaluate the population convergence state and diversity level, the local position distribution entropy is used to dynamically adjust the individual angle disturbance factor β (t) and jump step factor J (t) ; S36. According to the jump update strategy, in the current iteration round, a jump update is performed based on the individual's current parameter vector, the optimal individual position and direction angle: in, is the updated parameter combination vector of the kth individual in the t+1th round of iteration, is the current parameter combination vector of the kth horned lizard individual in the tth iteration, J (t) is the jump step coefficient in the tth iteration, cos is the cosine function, is the individual parameter combination with the highest fitness score in the tth iteration; S37. In each iteration, each updated parameter combination is applied to the OCR recognition module and the ASR recognition module respectively, and the extracted image frame sequence and audio stream are recognized and processed to obtain the image text information set respectively. Voice and text information collection S38. Based on the recognition results of the image and speech, calculate the recognition accuracy, processing delay, resource consumption, modal confidence consistency and confidence propagation edge stability of the individual parameter combination in the current round, and combine the above eigenvalues into a eigenvector Z (k) , input into the joint evaluation network to obtain the individual fitness score F (k) ; S39, perform iterative fitness evaluation on each individual, compare the fitness score of this round with the score of the previous round, if F (k,t+1) >F (k,t) , then keep the new parameter combination and update the direction memory vector Otherwise, keep the original parameters unchanged and update the individual direction disturbance state, where F (k,t+1) In the horned lizard optimization algorithm, the fitness score of the kth individual in the t+1th iteration, F (k,t) represents the fitness score of the same individual in the tth iteration; S310, repeat steps S33 to S39 until the set maximum number of iterations T is reached max Or the global optimal fitness change satisfies the convergence judgment threshold ε, and finally outputs the parameter combination X with the optimal fitness score * .
5. The method for collecting long video content information based on OCR and speech recognition technology according to claim 4 is characterized in that: The S4 specifically includes: S41, the optimal parameter combination vector output by the horned lizard optimization algorithm They are respectively configured into the OCR recognition module and the ASR recognition module as the final recognition parameters. The OCR recognition module is applied to the video image frame sequence, and the ASR recognition module is applied to the audio stream data, where Δf * is the optimized image frame frequency difference, is the optimized image clarity threshold, is the optimized text confidence threshold is the minimum size threshold of the optimized text candidate area, is the minimum length of the optimized speech slice window, α * is the optimized audio semantic segment correlation coefficient threshold; S42, for all the segmented video contents obtained by preprocessing, call the OCR recognition module and the ASR recognition module segment by segment according to the segment number i, and use the optimal parameter combination vector X * Perform identification processing; S43, in each segment i, the OCR recognition module uses parameters Process the image frame sequence, extract the text region candidate box, confidence score and image character results, and obtain the image text information set; S44, in each segment i, the ASR recognition module uses parameters Slice the audio stream and extract voiceprint features, and perform speech-to-text conversion to obtain a speech-to-text information set; S45, the image text information set obtained for each segment i Voice and text information collection Perform preliminary alignment based on timestamps. The alignment strategy matches the appearance frame position of the text segment with the time window overlap of the speech slice to form a segmented fusion text set. S46, merging the fused text sets of all video segments i to obtain the overall optimized recognized image text information and voice text information; S47. Send the merged image text information and speech text information to the belief propagation processing module respectively to perform semantic reasoning, fusion and final information structure generation.
6. The method for collecting long video content information based on OCR and speech recognition technology according to claim 1, characterized in that: The S5 specifically includes: S51, based on the acquired image text information set T img and the voice text information set T speech , construct the modality fusion factor graph structure G = (V, F, E), where V represents all modality semantic variable nodes to be jointly reasoned, including the image segment text variable node v i With the speech segment text variable node v j , F represents the set of semantic relationship factors between modalities, and E represents the set of edges connecting semantic variable nodes and relationship factors; S52, constructing a set of modal factor functions Φ={φ ij }, used to model the semantic relevance scoring function between image text segments and speech text segments, defined as: Among them, Emb(v i ) represents the text fragment v i FastText embedding vector, σ is the smoothing factor that adjusts the influence range of the embedding space distance, Emb(v j ) represents the text fragment v j FastText embedding vector, exp is the exponential function; S53, for each variable node v k Initialize the prior marginal probability distribution Normalized generation is performed based on the confidence score in the OCR recognition module or the ASR recognition module; S54, for the constructed factor graph G, the variational belief propagation algorithm is used to iteratively calculate the edge probability, and for each connecting edge e ij , execute the message passing rules: in, represents the message transmitted from node i to node j in round t+1, N(i) represents the set of adjacent nodes of node i, represents the message transmitted from node i to node j in round t; S55. After each round of iteration, the edge probabilities of all variable nodes are updated using a normalization correction strategy: Among them, Z j is the normalization factor, Indicates that in the t+1th iteration, the variable node v in the factor graph j The marginal probability distribution function of S56, repeat steps S54 to S55 until the edge probabilities of all variable nodes converge or the maximum number of iterations T is reached. BP , at this time the final marginal probability distribution As the joint semantic confidence output of the text fragments, according to the joint semantic confidence output, the text fragments in the image text and the speech text with confidence higher than the set threshold θ are marked as semantic blocks to form a multimodal semantic block set.
Citation Information
Cited By
Electric meter reading method, device and equipment based on image recognition and medium
CN120894771A
Intelligent AI information processing method and system
CN121354118A