Multi-modal visual analysis system for teaching
Through the multimodal visual analysis system, the real-time, adaptability and professional feature recognition problems in electronic circuit practical teaching are solved, and all-round practical guidance is achieved, which improves teaching efficiency and quality.
Patent Information
- Application Number
- CN202510828568.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing technology has problems such as insufficient real-time, poor adaptability, single functions, and low professional feature recognition accuracy in practical teaching of electronic circuits, and it is difficult to provide comprehensive practical guidance.
Design a multimodal visual analysis system for teaching, including video acquisition and processing module, multi-threaded parallel analysis module, user interface module, system stability and resource management module, and log and error processing module. Through dynamic frame difference detection, multi-threaded parallel processing, interface layout optimization and resource monitoring, real-time video analysis and all-round guidance are realized.
It improves the practical efficiency and quality of electronic circuit practical teaching, provides comprehensive guidance on behavior description, error correction, operation classification and knowledge point push, and meets professional requirements.
Smart Images

Figure CN120339924A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly to a multi-modal visual analysis system for teaching. Background Art
[0002] In the field of practical teaching of electronic circuits, traditional guiding technologies face many bottlenecks. From the perspective of real-time performance, the guiding method based on manual observation is not only inefficient, but also prone to missed error judgments due to subjective factors such as observer fatigue and distraction of attention, and the feedback has obvious delays; while off-line video analysis needs to be carried out after the operation is completed, and it is impossible to give guidance in time during the actual operation, making it difficult to meet the need for real-time error correction in the practical operation of electronic circuits.
[0003] In terms of intelligence, existing systems mostly adopt pattern matching algorithms based on preset rules. For example, a fixed threshold of the solder joint shape is set to judge the welding quality. When special situations such as complex lighting conditions and slight component offsets occur, the system is difficult to accurately identify. This mechanism lacking adaptive ability is often helpless in the face of diverse operation scenarios and personalized operation habits.
[0004] In terms of functional integrity, most tools have relatively single functions. Taking the error detection tool as an example, although it can find errors in the operation, it cannot provide specific correction steps and relevant knowledge points for the errors; the operation classification tool can only classify the operation behaviors, but cannot push matching exercises to learners to consolidate knowledge, making it difficult to achieve effective guidance for the entire process of practical teaching of electronic circuits.
[0005] In terms of domain adaptability, the feature extraction methods adopted by general visual models, such as the general convolution kernels of convolutional neural networks, are difficult to accurately capture the subtle features of professional tools and components in electronic circuits. For professional information such as the color ring coding of resistors, the pointer scale of multimeters, the roundness and gloss of solder joints, and the capacitance polarity markings, the recognition accuracy of general models is relatively low, which cannot meet the professional requirements of practical teaching of electronic circuits.
[0006] Therefore, it is necessary to design a multi-modal visual analysis module for teaching to intelligently analyze the real-time video stream collected by the camera in the laboratory, help users complete key operations such as circuit soldering, component installation, and circuit debugging, and provide all-round guidance such as behavior description, error correction, operation classification, knowledge point push, and exercise generation, so as to improve the practical operation efficiency and quality. Summary of the Invention
[0007] The purpose of the present invention is to provide a multi-modal visual analysis system for teaching to solve the problems existing in the above background art.
[0008] To achieve the above object, the present invention provides a multi-modal visual analysis system for teaching, including: A video acquisition and processing module that real-time acquires video frames and calculates the difference between two adjacent frames of images through a dynamic frame difference detection algorithm to determine whether there is a dynamic change in the scene; A multi-thread parallel analysis module, including thread design and cooperation and a dynamic load balancing mechanism, to achieve parallel processing of video acquisition and analysis; A user interface module that displays the captured real-time video frames by designing the interface layout; An analysis result output and display module that real-time pushes and displays the intermediate results analyzed by the multi-thread parallel analysis module, and sends the final analysis conclusion after analysis to the user interface module; A system stability and resource management module, including multi-thread management and synchronization and resource overhead monitoring and optimization; A log and error handling module that records the running information of the program, captures exceptions at the main entry of the program, records them in the log and takes corresponding measures.
[0009] Preferably, for video acquisition, cv2.VideoCapture is used to open the camera, and video frames are continuously read in a loop. A buffer mechanism is adopted to cache the acquired video frames. The buffer serves as a temporary storage area. When the camera captures a video frame, it is first stored in the buffer and waits for subsequent processing. This can avoid frame loss caused by short-term network fluctuations or hardware failures; at the same time, by setting an appropriate frame rate (such as adding time.sleep(0.1)), the frame rate is controlled to avoid overly frequent processing and reduce the system burden. In practical applications, the frame rate can be dynamically adjusted according to the performance and requirements of the system.
[0010] Preferably, the dynamic frame difference detection algorithm determines whether there is a dynamic change in the scene by analyzing the difference between two adjacent frames of images. When the difference exceeds the threshold, it is considered that a dynamic change has occurred in the scene, and the current frame is analyzed. Specifically: Initialization: Initialize a variable to store the previous frame image, and set the initial value to None; this variable will be used in subsequent frame difference calculations; Loop to read frame data: Continuously read frame data in an infinite loop, and obtain the current frame image in each loop; Calculate the frame difference: ; Among them, represents the pixel value of the difference image at position ; and respectively represent the pixel values of the current frame and the previous frame image at position ; Grayscale processing: Convert the obtained difference image into a grayscale image; the purpose of grayscale processing is to simplify subsequent processing steps. A grayscale image has only one channel and is more convenient to process.
[0011] Binarization processing: Perform binarization processing on the grayscale image. Binarization divides the pixel values in the image into two categories according to a preset threshold. Pixels greater than the threshold are set to 255 (white), and pixels less than or equal to the threshold are set to 0 (black), expressed as: ; where, represents the pixel value of the binarized difference image at position ; represents the pixel value of the grayscale difference image at position ; is the preset threshold, set to 10% here; Calculate the difference percentage: ; where, is the number of non-zero pixels; is the total number of pixels; Judge whether to analyze: Compare the calculated difference percentage with the preset threshold. If the difference percentage exceeds the threshold, it is considered that the change between two frames is large enough and the current frame needs to be further analyzed. In the code, the current frame will be encoded as byte data in JPEG format and sent out through a signal; Update the previous frame: Assign the current frame to the variable storing the previous frame image for use as the previous frame in the next loop; Processing of the first frame: When the variable storing the previous frame image is None, it indicates that the current is the first frame. For the first frame, it is encoded as byte data in JPEG format and sent out through a signal for analysis.
[0012] Preferably, the thread design and cooperation are achieved by designing two thread classes, CameraThread and AnalysisWorker. CameraThread is used for video frame acquisition, continuously reads video frames from the camera, and passes them to the AnalysisWorker thread for analysis. Specifically: Video frame decoding: Convert the video frame received by AnalysisWorker in the form of byte data (frame_data) into a numpy array; use the imdecode function of OpenCV to decode the numpy array into an image frame; Feature extraction: Use OpenCV's resize function to reduce the resolution of the image frame to a specified size; convert the processed image frame to a PIL image object; save the PIL image object as a byte stream in JPEG format and perform base64 encoding; Model inference: Call the ollama.generate function, pass in the model name, prompt information, and base64 encoding parameters of the image, and start streaming inference; process each output block in the inference process and check whether it has timed out. If so, throw a timeout exception; splice each output block into a complete analysis result.
[0013] Through multithreading, parallel processing of video acquisition and analysis is achieved to improve the overall performance of the system. The pyqtSignal signal mechanism is used to realize communication between threads and return the analysis results to the main thread for display in a timely manner.
[0014] Preferably, the dynamic load balancing mechanism adopts a dual mechanism of timed forced analysis and frame difference triggering to ensure real-time performance; the timed forced analysis analyzes the current video frame at a fixed time interval (10 seconds), and will trigger the analysis operation even when there is no obvious change in the video picture, thereby ensuring that the system can continuously monitor the video content and avoid missing important information due to the limitations of frame difference detection; the dual mechanism of frame difference triggering is consistent with the calculation method of the dynamic frame difference detection algorithm. In actual applications, the time interval and frame difference threshold of the timed forced analysis can be adjusted according to different scenarios and needs.
[0015] Preferably, the interface layout design uses PyQt5 to build the user interface, and the interface layout is performed through layout managers such as QVBoxLayout and QHBoxLayout to ensure that the interface elements are neatly arranged and beautiful. Reasonable division of interface areas, such as video display area, control button area and analysis result display area, is convenient for users to operate and view information. When designing the interface layout, it is necessary to consider the user's usage habits and operational convenience to avoid the interface being too complicated or confusing; the captured video frame is converted into QPixmap format through video display and interactive optimization, and displayed on QLabel to achieve real-time display of the video. In order to improve the display effect, some preprocessing can be performed on the video frame, such as adjusting brightness and contrast. Add corresponding slot functions to the button, and when the user clicks the button, the corresponding operation can be triggered in time, such as starting / stopping the camera, forcing analysis, etc. At the same time, optimize the rendering performance of the interface and reduce the interface jamming phenomenon. Double buffering technology can be used to prepare the image of the next frame in advance to reduce rendering time.
[0016] Preferably, the analysis result output and display module adopts a dual-signal output mechanism, and uses two different signals to push the intermediate result and the final analysis conclusion in real time respectively; for the intermediate result signal, during the model inference process of the AnalysisWorker thread, phased intermediate results will be continuously generated. Every time an intermediate result block is obtained, it will be sent out through a specific pyqtSignal. This signal carries two key pieces of information, one is the text box index corresponding to the analysis result, and the other is the specific intermediate result content; in this way, the user interface can receive these intermediate results in real time and update and display them, allowing users to track the progress of the analysis in a timely manner. The final analysis conclusion is the definitive result obtained by the AnalysisWorker thread after comprehensively and systematically analyzing the video frames. It is obtained after completing a series of complex operations (such as decoding video frames, extracting key features, using the model for inference, etc.). During the entire analysis process, the model will comprehensively consider various information in the video frames and obtain the final conclusion through calculation and judgment. This conclusion is presented in the form of a string, containing a detailed interpretation, analysis, and judgment of the video frame content, such as the operation steps, operation types, potential errors or risks in the picture, as well as the questions and answers of relevant knowledge points. Finally, this conclusion will be sent to the user interface through a specific signal and displayed completely, enabling users to clearly understand the key information and analysis results contained in the video frame. The final analysis conclusion signal sends this final analysis conclusion through another pyqtSignal. Similarly, this signal also carries the text box index and the complete analysis result.
[0017] Result display optimization: Format the analysis results and display them to users in a clear and understandable way. Various forms such as text and charts can be used for display to improve the readability of the results. At the same time, provide detailed explanations and descriptions of the results to help users better understand the analysis results.
[0018] Preferably, in multi-thread management and synchronization, a lock mechanism and semaphores are used to control the access of threads to shared resources, avoiding resource competition and deadlock problems; for example, when operating on shared video frame data, a mutex lock is used to ensure that only one thread can access this data at the same time. At the same time, reasonably design the life cycle of threads and destroy the threads in a timely manner when they are not needed to release system resources. Thread pool technology can be adopted to uniformly manage and schedule threads, improving the reusability and efficiency of threads. The lock mechanism is used to ensure that there is one thread accessing the shared video frame data at the same moment, guaranteeing the accuracy and consistency of the data; the semaphore is used to limit the number of simultaneously running AnalysisWorker threads. At the same time, the use of the buffer is controlled by controlling the access to the queue.
[0019] Resource Overhead Monitoring and Optimization: Regularly monitor the resource usage of the system, such as CPU usage rate, memory occupancy, etc. When the resource usage exceeds a certain threshold, take corresponding optimization measures, such as reducing unnecessary computing tasks, releasing cached data, etc. At the same time, optimize algorithms and data structures to improve the execution efficiency of the code and reduce resource overhead. For example, use more efficient algorithms for image feature extraction to reduce the amount of calculation and memory occupancy.
[0020] Preferably, the logging and error handling module configures logs using the logging module, records the running information of the program into the app.log file for logging. Through logging, it is convenient to troubleshoot problems that occur during the program running process; by setting different log levels, different levels of detailed information can be recorded according to requirements. For example, during the development and debugging stage, the log level can be set to DEBUG to record more detailed information; after the official launch, the log level can be set to INFO to only record key information; use the try-except statement at the main entry of the program to catch exceptions. When serious errors occur in the program, record the error information in the log and pop up a message box to prompt the user. Classify and handle different types of exceptions, and take corresponding recovery measures, such as restarting threads, releasing resources, etc., to improve the fault tolerance of the system. For example, when the camera fails to open, prompt the user to check the camera device and try to open it again.
[0021] Therefore, by adopting the above-mentioned multimodal visual analysis system for teaching, the present invention can intelligently analyze the real-time video stream collected by the camera in the laboratory, help users complete key operations such as circuit soldering, component installation, and circuit debugging, provide all-round guidance such as behavior description, error correction, operation classification, knowledge point push, and exercise generation, and improve the practical operation efficiency and quality.
[0022] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0023] Figure 1 It is a schematic structural diagram of a multimodal visual analysis system for teaching according to the present invention. Detailed Embodiments
[0024] The following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0025] Please refer to Figure 1, A multi-modal visual analysis system for teaching, including: I. Initialization Core function: Complete the environment configuration, resource loading, and global parameter initialization before the program starts.
[0026] Key components: 1. Log configuration Use the logging module to write logs to app.log, recording the system running status and error information (such as camera connection failure, model analysis timeout).
[0027] Log level: INFO and above, and the format includes timestamp, module name, log level, and message.
[0028] 2. Qt environment configuration
[0029] Enable high DPI adaptation , ensuring that the interface is displayed normally on high-resolution screen devices.
[0030] Set the Fusion style and custom color palette to define the visual style (color, border, font) of interface elements (buttons, text boxes, progress bars).
[0031] 3. Memory optimization (Windows platform)
[0032] Call the system function SetProcessWorkingSetSize through ctypes to limit the memory usage range of the program and avoid performance problems caused by memory leaks.
[0033] 4. Interaction logic: The program entry (if __name__ == '__main__') executes the initialization logic first to ensure that subsequent modules run in a unified configuration environment.
[0034] II. Interface interaction
[0035] Core function: Build a user operation interface, process user input (such as starting / stopping the camera), and display real-time video and analysis results.
[0036] Key components: 1. Main window layout (RealTimeAnalysisTool) Left video panel: The QFrame area displays the real-time camera image (rendered by QLabel) and supports adaptive scaling. Control buttons (start / stop the camera) and status indicators trigger the start and stop of the camera thread through click events.
[0037] Right analysis panel: The QGridLayout generates multiple "analysis cards", and each card contains: An analysis title (such as "Real-time Operation Analysis"), a description (analysis prompt words), a progress bar (QProgressBar), and a result text box (QTextEdit, read-only).
[0038] The QScrollArea ensures that when the analysis results exceed the interface height, they can be scrolled for viewing.
[0039] 2. Status bar: Displays the system status (such as "Camera is running") and the current time (updated every second).
[0040] User interaction logic
[0041] 1. Camera control: Click the button to switch the camera status, and use the toggle_camera method to switch the button text, style, and the color of the status indicator.
[0042] Analysis result display: Through the signal-slot mechanism, receive the streaming output (analysis_stream) and the final result (analysis_complete) of the analysis thread, and update the content of the text box and the status of the progress bar.
[0043] 2. Data flow
[0044] The new_display_frame signal of the camera thread → the _update_video_frame slot function → decode the frame data and render it to the QLabel.
[0045] The analysis_stream signal of the analysis thread → the _update_analysis_stream slot function → append the analysis results to the text box line by line.
[0046] The analysis_complete signal of the analysis thread → the _finalize_analysis_output slot function → mark the analysis as completed and update the progress bar to 100%.
[0047] III. Processing
[0048] Core function: Real-time collection of the industrial camera video stream, preprocessing of the frame data (frame difference detection), and sending valid frames (display frames / analysis frames) through signals.
[0049] Key components: 1. Camera thread class (CameraThread) Initialization: Supports multiple backends (DSHOW / MSMF / ANY), and preferentially uses DSHOW to ensure Windows compatibility.
[0050] 2. Frame buffer queue (deque, maximum length 5) to avoid memory overflow, and a retry mechanism (up to 3 times) to handle camera read failures.
[0051] 3. Core logic (run method): Frame capture: Continuously read camera frames, with a fixed resolution of 640x480, encoded in JPEG format (quality 70% to reduce network transmission pressure).
[0052] Frame difference detection: Calculate the pixel difference between the current frame and the previous frame. When the difference exceeds 10%, mark it as a "frame to be analyzed" (to avoid meaningless repeated analysis).
[0053] 4. Signal sending: new_display_frame: Send all valid frames for real-time display.
[0054] new_analysis_frame: Only send frames with significant changes or the first frame (to trigger the analysis logic).
[0055] 5. Exception handling: Catch errors such as camera initialization failure and read timeout, and send an empty frame signal to notify the interface to display the error status.
[0056] 6. Interaction logic
[0057] The main thread controls the lifecycle of the camera thread through the start() / stop() methods.
[0058] Frame data is transmitted through a binary byte stream (bytes) to avoid memory conflicts caused by directly operating on OpenCV arrays across threads.
[0059] IV. Intelligent analysis
[0060] Core functions: Preprocess the analysis frames sent by the camera, call the Ollama model for multimodal reasoning, and return streaming analysis results, specifically including: 1. Use multithreading to achieve parallel analysis; 2. Image preprocessing: Resize (160×120) and Base64 encode; 3. Streaming API call and result processing; 4. Timeout control (20 seconds) and automatic retry (up to 3 times).
[0061] The implementation code is as follows: class AnalysisWorker(QThread): def attempt_analysis(self): """Attempt to perform analysis""" # Image preprocessing buffer = np.frombuffer(self.frame_data, dtype=np.uint8) frame = cv2.imdecode(buffer, cv2.IMREAD_COLOR) frame = cv2.resize(frame, self.frame_resize) base64_img = self._frame_to_base64(frame) # Construct the request request_data = { "model": self.ollama_model, "prompt": self.prompt, "stream": True, "images": [base64_img], "options": {"temperature": 0.7, "num_predict": 1024} } # Send the request and handle the streaming response full_response = "" response = requests.post( f"{self.ollama_server} / api / generate", json=request_data, stream=True, timeout=30 ) for line in response.iter_lines(): if self.should_stop: break chunk = json.loads(line.decode('utf-8')) if not chunk.get('done', False): response_chunk = chunk.get('response', '') full_response += response_chunk self.analysis_stream.emit(self.box_index, response_chunk) if not self.should_stop: self.analysis_complete.emit(self.box_index, full_response)
[0062] Key components: 1. AnalysisWorker class Input parameters: Frame data (binary byte stream), analysis prompt (such as "Detect possible errors in the current operation"), target text box index (corresponding to the right analysis card).
[0063] Preprocessing logic: Decode the frame data into an OpenCV array and scale it to 160x120 (reduce the model input resolution and improve the inference speed).
[0064] Convert it to RGB format and encode it as a base64 string through PIL.Image (to conform to the image input format of the Ollama model).
[0065] 2. Model call: Use the ollama.generate interface, specify the model modelscope.cn / lmstudio-community / MiniCPM-o-2_6-gguf:latest, and enable streaming output (return results word by word).
[0066] 3. Timeout protection: The maximum analysis time is 30 seconds per single analysis. If it exceeds, retry (up to 3 times).
[0067] 4. Signal mechanism: analysis_start: Sent when the analysis starts, triggering the interface to display the timestamp and separator line.
[0068] analysis_stream: Return the model output chunk by chunk, and update the text box content in real time.
[0069] analysis_complete: Sends final results when analysis is complete (success / failure).
[0070] 5. Performance Optimization
[0071] Image compression: The analysis frame quality is set to 70% to reduce the amount of data transmitted over the network and processed by the model.
[0072] Parallel processing: 4 analysis tasks (corresponding to 4 text boxes) run in parallel through independent threads to improve throughput.
[0073] 5. Thread Management
[0074] Core functions: Coordinate the life cycles of the camera thread and the analysis thread to ensure safe multi-threaded interaction; trigger forced analysis at regular intervals to avoid missed detections.
[0075] Key components: Thread safety mechanism: Signal-slot communication: All cross-thread data interactions (such as frame data transmission and analysis result updates) are implemented through PyQt's signal slots, avoiding thread safety issues caused by direct operation of UI components.
[0076] Worker thread list (analysis_workers): records all running analysis threads, terminates and cleans them up in batches when the camera stops to avoid memory leaks.
[0077] Timed forced analysis
[0078] The timer (10-second interval) triggers the _force_analyze_frame method, forcing the use of the latest frame for analysis regardless of whether the frame difference meets the standard.
[0079] Application scenario: Prevent analysis stagnation when there is no significant image change for a long time (such as workers continuing the same operation).
[0080] Exception handling
[0081] When the camera thread fails to read, it will automatically retry to open the device (within 3 times). If the number of times exceeds this, the interface will prompt an error and stop the thread.
[0082] When the analysis thread times out or the model call fails, a retry mechanism (3 times) and error message feedback (such as "Analysis timed out, retry failed") are implemented.
[0083] 6. Auxiliary tools
[0084] Core functions: Provide general tool functions, exception capture and resource release logic to enhance system stability.
[0085] Key components: 1. Image Tools _frame_to_base64(AnalysisWorker): Convert the OpenCV frame to a base64 string for recognition by the Ollama model.
[0086] cv2.imencode / cv2.imdecode: Convert frame data between binary byte streams and OpenCV arrays.
[0087] 2. Resource Release
[0088] Stop the camera thread in the window close event (closeEvent), release the camera handle (cap.release()), and destroy the OpenCV window (cv2.destroyAllWindows()).
[0089] After the analysis thread is completed, trigger _remove_worker through the finished signal to remove the ended thread from the list.
[0090] 3. Logging and Error Messages
[0091] Each module records key events (such as "Camera thread started", "Analysis timeout") through an independent logger instance for easy debugging.
[0092] QMessageBox displays user-level errors (such as "Failed to start the camera"), combined with technical-level error records in the log (exc_info = True).
[0093] The workflow of a multi-modal visual analysis system for teaching in this embodiment is as follows: 1. System Initialization Main window loading: Create a PyQt5 main window and set the resolution (1360×860).
[0094] Initialize the left video panel (real-time camera feed) and the right analysis panel (4 analysis channels).
[0095] Configure the logging system (log to app.log).
[0096] Set Windows memory optimization (SetProcessWorkingSetSize).
[0097] Camera thread initialization: Try multiple backends (CAP_DSHOW>CAP_MSMF>CAP_ANY).
[0098] Set the resolution (640×480@30fps).
[0099] Initialize the frame difference detection mechanism (cv2.absdiff).
[0100] Analysis thread pool initialization: 4 independent threads, respectively processing: operation step analysis, operation type classification, potential error detection, and knowledge point question generation. Each thread is bound to an independent signal slot for streaming the analysis results.
[0101] 2. Video acquisition process
[0102] Camera startup: The user clicks "Start Camera", triggering CameraThread.start().
[0103] The thread enters the loop to read frames mode: while self.running: ret,frame=self.cap.read() If not ret: self.retry_count += 1 # Retry 3 times in case of failure Frame difference detection (triggered by dynamic analysis): Calculate the difference between the current frame and the previous frame: python diff=cv2.absdiff(prev_frame, current_frame) diff_gray = cv2.cvtColor(diff, cv2.COLOR_BGR2GRAY) _, diff_threshold = cv2.threshold(diff_gray, 30, 255, cv2.THRESH_BINARY) diff_percentage = (np.count_nonzero(diff_threshold) / diff_threshold.size) 100 If diff_percentage > 10%, trigger the analysis.
[0104] Dual-channel output: Display the frame (JPEG 70% quality) → Send it to the UI thread for rendering.
[0105] Analyze the frame (downsample to 160×120 + JPEG compression) → Send it to the AI analysis thread.
[0106] 3. AI Analysis Process
[0107] Analysis Task Distribution: Four independent analysis tasks (AnalysisWorker) are triggered per frame.
[0108] Each task is bound to a different prompt: python prompts = "Analyze the current electrical operation step", "Identify the operation type (wiring / measurement / debugging)", "Detect potential errors or risks", "Generate relevant knowledge point test questions" Model Invocation (Ollama API): Use the MiniCPM - o - 2_6 model (lightweight multi - modal model).
[0109] Image Encoded as Base64 for Transmission: python img = Image.fromarray(frame) buffer = io.BytesIO() img.save(buffer, format="JPEG", quality=70) base64_img = base64.b64encode(buffer.getvalue()).decode("utf - 8") Streaming Invocation (stream = True): python response = ollama.generate( model="MiniCPM - o - 2_6", prompt=prompt, images=[base64_img], stream=True, options={"timeout": 30} ) Streaming Result Update: Send analysis fragments to the UI thread in real - time: python for chunk in response: if not chunk['done']: self.analysis_stream.emit(box_index, chunk['response']) The final result is marked as " Analysis completed".
[0110] 4. UI Rendering and Interaction
[0111] Video display: Receive the new_display_frame signal and update the QPixmap: python q_img = QImage(rgb_data, w, h, bytes_per_line, QImage.Format_RGB888) pixmap = QPixmap.fromImage(q_img).scaled(video_label.size(), Qt.KeepAspectRatio) video_label.setPixmap(pixmap) Analysis result display: Progress bar animation (10% → 100%).
[0112] Streaming text rendering (with timestamps): Error handling: Camera error → Display "Camera error" (red status bar).
[0113] Analysis timeout → Display "Analysis timeout" (orange warning).
[0114] 5. System Shutdown
[0115] Resource release: Stop the camera thread (CameraThread.stop()).
[0116] Terminate all analysis threads (worker.stop()).
[0117] Release OpenCV resources (cv2.destroyAllWindows()).
[0118] Logging: Record the exit status (app.log).
[0119] Therefore, the present invention adopts the above-mentioned multi-modal visual analysis system for teaching, and improves the practical operation efficiency and quality by analyzing real-time video frames.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-modal visual analysis system for teaching, characterized in that, Including: A video acquisition and processing module that real-time acquires video frames and calculates the difference between two adjacent frames of images through a dynamic frame difference detection algorithm to determine whether there are dynamic changes in the scene; A multi-threaded parallel analysis module that includes thread design and collaboration and a dynamic load balancing mechanism to achieve parallel processing of video acquisition and analysis; A user interface module that displays the captured real-time video frames by designing the interface layout; An analysis result output and display module that real-time pushes and displays the intermediate results analyzed by the multi-threaded parallel analysis module and sends the final analysis conclusion after analysis to the user interface module; A system stability and resource management module that includes multi-thread management and synchronization and resource overhead monitoring and optimization; A log and error handling module that records the running information of the program, captures exceptions at the main entry of the program, records them in the log and takes corresponding measures.
2. The multimodal visual analysis system for teaching according to claim 1, wherein: For video acquisition, cv2.VideoCapture is used to open the camera, continuously reads video frames in a loop, adopts a buffer mechanism to cache the acquired video frames. The buffer serves as a temporary storage area waiting for subsequent processing; at the same time, time.sleep is added to control the frame rate.
3. A multimodal visual analysis system for teaching according to claim 1, characterized in that, The dynamic frame difference detection algorithm determines whether there are dynamic changes in the scene by analyzing the difference between two adjacent frames of images. When the difference exceeds the threshold, it is considered that there are dynamic changes in the scene, and then the current frame is analyzed. Specifically: Initialization: Initialize a variable to store the previous frame of image, and set the initial value to None; Loop to read frame data: Continuously read frame data in an infinite loop, and each loop will obtain the current frame of image; Calculate the frame difference: ; Among them, represents the pixel value of the difference image at the position ; and respectively represent the pixel values of the current frame and the previous frame image at the position ; Grayscale processing: Convert the obtained difference image to a grayscale image; Binarization processing: Perform binarization processing on the grayscale image, expressed as: ; Among them, represents the pixel value of the binarized difference image at the position ; represents the pixel value of the grayscale difference image at the position ; is a preset threshold value; Calculate the difference percentage: ; Among them, is the number of non-zero pixels; is the total number of pixels; Judge whether to analyze: Compare the calculated difference percentage with a preset threshold. If it exceeds the threshold, then analyze; Update the previous frame: Assign the current frame to the variable storing the previous frame of image for use as the previous frame in the next loop; Processing of the first frame: When the variable storing the previous frame of image is None, it indicates that the current is the first frame. For the first frame, encode it into byte data in JPEG format and send it out through a signal for analysis.
4. A multimodal visual analysis system for teaching according to claim 1, wherein Thread design and collaboration are achieved by designing two thread classes, CameraThread and AnalysisWorker. CameraThread is used for video frame acquisition, continuously reads video frames from the camera and passes them to the AnalysisWorker thread for analysis. Specifically: Video frame decoding: Convert the video frame received by AnalysisWorker in the form of byte data into a numpy array; use the imdecode function of OpenCV to decode the numpy array into an image frame; Feature extraction: Use the resize function of OpenCV to reduce the resolution of the image frame to the specified size; convert the processed image frame into a PIL image object; save the PIL image object as a JPEG format byte stream and perform base64 encoding; Model inference: Call the ollama.generate function, pass in the model name, prompt information, and the base64 encoding parameter of the image, and start streaming inference; process each output block during the inference process, check for timeouts, and if a timeout occurs, throw a timeout exception; Concatenate each output block into a complete analysis result.
5. A multimodal visual analysis system for teaching according to claim 1, characterized in that: The dynamic load balancing mechanism adopts a dual mechanism of timed forced analysis and frame difference trigger; the timed forced analysis analyzes the current video frame at fixed time intervals; the frame difference trigger dual mechanism is consistent with the calculation method of the dynamic frame difference detection algorithm.
6. The multimodal visual analysis system for teaching according to claim 1, wherein: The interface layout design uses PyQt5 to build the user interface, performs interface layout through the layout manager, and reasonably divides the interface area; converts the captured video frame into QPixmap format through video display and interaction optimization and displays it on QLabel to achieve real-time video display.
7. A multimodal visual analysis system for teaching according to claim 4, characterized in that The analysis result output and display module adopts a dual-signal output mechanism, and uses two different signals to push the intermediate result and the final analysis conclusion in real time respectively; the intermediate result signal is that during the model inference process of the AnalysisWorker thread, it will continuously generate phased intermediate results. Every time an intermediate result block is obtained, it will be sent out through pyqtSignal. This signal carries two key pieces of information, one is the text box index corresponding to the analysis result, and the other is the specific intermediate result content; the final analysis conclusion is the result obtained by the AnalysisWorker thread analyzing the video frame. The final analysis conclusion signal sends this final analysis conclusion through another pyqtSignal. Similarly, this signal also carries the text box index and the complete analysis result.
8. A multimodal visual analysis system for teaching according to claim 7, characterized in that: In multi-thread management and synchronization, a lock mechanism and semaphore are used to control thread access to shared resources; the lock mechanism is used to ensure that there is one thread accessing the shared video frame data at the same time; the semaphore is used to limit the number of simultaneously running AnalysisWorker threads. At the same time, the use of the buffer is controlled by controlling access to the queue.
9. A multimodal visual analysis system for teaching according to claim 1, wherein: The log and error handling module configures the log using the logging module, records the running information of the program into the app.log file for logging, and by setting different log levels, different levels of detailed information can be recorded according to requirements; at the main entry of the program, use try-except statements to catch exceptions. When a serious error occurs in the program, the error information is recorded in the log and a message box is popped up to prompt the user.
Citation Information
Patent Citations
Safety monitoring system and method based on cloud computing
CN116980569A
Table data interactive processing method based on large language model
CN118394909A
To-be-detected image screening method, device and equipment and computer readable storage medium
CN119271419A
Multi-service bearing teaching video analysis method and system based on cross-modal data
CN120014510A
Automatic learning engine device based on packaging large model training platform
CN120066523A
Cited By
Vehicle-mounted video delay elimination method and system based on multi-thread optimization
CN121000922A