Video picture processing terminal and video picture processing method
By extracting visual and audio features through a video processing terminal, enhanced video images are generated, solving the problem that speakers cannot operate and explain simultaneously in remote meetings or teaching, and achieving more efficient communication and interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LANTO ELECTRONIC LIMITED
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-05
AI Technical Summary
In remote meetings or remote teaching, speakers cannot focus on inputting information and explaining simultaneously when sharing content, resulting in insufficient interactivity and clarity of instructions, making it difficult for remote participants to accurately interpret the content being instructed.
The video processing terminal extracts visual, auditory, and audio features to generate enhanced video images. Combined with the speaker's physical actions and audio content, this improves interactivity and clarity of instructions.
Enhanced video displays enable remote participants to clearly understand the speaker's key points, improving communication efficiency and the accuracy of information delivery, and increasing interactivity and clarity of instructions between the speaker and remote participants.
Smart Images

Figure CN121985153A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing, and more particularly to a video processing terminal and a video processing method. Background Technology
[0002] Currently, remote conferencing or remote teaching often involves sharing screens or presentations, supplemented by recording live footage of the speaker so that remote participants can access the presentation content.
[0003] Generally, speakers use a mouse or other input tools to add markers or indicators to the content being shared (such as presentations or documents) to prompt remote participants. However, controlling the input tool during the presentation prevents the speaker from focusing on the content, which is clearly inconvenient for them.
[0004] Furthermore, in live projection scenarios, speakers can use laser pointers or pointers to indicate projected content. However, remote participants can only roughly determine the location the speaker is pointing to from the live image, and cannot accurately determine the content being indicated.
[0005] Therefore, how to propose a terminal and method that can increase interactivity and clarity of instructions is one of the problems that this field seeks to solve. Summary of the Invention
[0006] This application provides a video processing terminal, which includes an image processing module, a visual processing module, a voice processing module, a decision processing module, and a screen processing module. The image processing module extracts screen feature information from image screen information, the screen feature information including text information and object information. The visual processing module extracts visual feature information from real-time local images. The voice processing module extracts voice feature information from real-time voice information. The decision processing module determines the update content of the basic image screen based on the screen feature information, the visual feature information, and the voice feature information. The screen processing module generates an enhanced image screen based on the updated content and the basic image screen, and the enhanced image screen is transmitted as a video screen to an external remote device.
[0007] This application provides a video processing method, comprising: extracting image feature information from image information, the image feature information including text information and object information; extracting visual feature information from real-time local images; and extracting voice feature information from real-time voice information; determining update content of a basic image frame based on the image feature information, the visual feature information, and the voice feature information; and generating an enhanced image frame based on the updated content and the basic image frame, the enhanced image frame being used as a video frame for transmission to an external remote device.
[0008] This application provides a video processing terminal, which includes a memory, an image capturing device, and a microprocessor. The memory stores a program for a video processing system. The image capturing device is used to obtain real-time local images. The microprocessor is coupled to the memory and the image capturing device to execute the program of the video processing system to implement a corresponding video processing method. The video processing method includes: extracting image feature information from image image information, the image feature information including text information and object information; extracting visual feature information from the real-time local images; and extracting voice feature information from real-time voice information; determining the update content of a basic image frame based on the image feature information, the visual feature information, and the voice feature information; and generating an enhanced image frame based on the updated content and the basic image frame, the enhanced image frame being used as a video frame to be transmitted to an external remote device.
[0009] This application embodiment determines the update content of the basic image frame based on the image feature information, the visual feature information, and the voice feature information, and automatically generates an enhanced image frame based on the speaker's physical operation and / or narration content according to the updated content and the basic image frame. This allows remote participants to clearly understand the speaker's key points through the enhanced image frame, improving communication efficiency and information transmission accuracy, effectively increasing the interactivity between the speaker and remote participants, and enhancing the clarity of the speaker's instructions. Attached Figure Description
[0010] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0011] Figure 1 This is a schematic diagram of the application environment of a video processing terminal according to an embodiment of this application;
[0012] Figure 2 This is a schematic diagram of another application environment of the video processing terminal according to an embodiment of this application;
[0013] Figure 3 This is a schematic diagram of the system architecture of a video processing terminal according to an embodiment of this application;
[0014] Figure 4 This is a schematic diagram of another system architecture for a video processing terminal according to an embodiment of this application;
[0015] Figure 5 This is a schematic diagram of the architecture of a video processing system according to an embodiment of this application;
[0016] Figure 6 This is a schematic diagram illustrating a flowchart of a video processing method according to an embodiment of this application;
[0017] Figure 7 This is a schematic diagram of another embodiment of the video image processing method according to the present application;
[0018] Figure 8 This is a schematic diagram of yet another embodiment of the video image processing method according to the present application;
[0019] Figure 9 This is a schematic diagram of another embodiment of the video image processing method according to the embodiments of this application;
[0020] Figure 10 This is a schematic diagram of another embodiment of the video image processing method according to the present application;
[0021] Figure 11 This is a schematic diagram of another embodiment of the video image processing method according to the present application.
[0022] Explanation of reference numerals in the attached diagram: 1. Video processing terminal; 2. Electronic device; 3. External remote device; 4. Projection device; 110. Processing module; 120. Storage module; 130. Communication module; 140. Image acquisition module; 150. Audio receiving module; 160. Projection module; 191. Image processing module; 192. Visual processing module; 193. Voice processing module; 194. Decision processing module; 195. Image processing module. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] Please refer to Figure 1 , Figure 1This is a schematic diagram of an application environment for a video processing terminal according to an embodiment of this application. The video processing terminal 1 is coupled to an electronic device 2. The electronic device 2 is coupled to an external remote device 3 via a network. The electronic device 2 is, for example, a smartphone, tablet, or laptop computer, and this application is not limited thereto. The video processing terminal 1 is used to capture real-time local images including those of the speaker, identify the speaker's physical operations, and generate corresponding enhanced video images or control commands based on the speaker's physical operations. The real-time local images include the speaker's physical operations and basic video images. Physical operations include, for example, the speaker's posture or gestures, the movement of a pointer held by the speaker, etc. In one embodiment, the basic video image is, for example, the content of a presentation or document being projected. The basic video image can be projected by the video processing terminal 1 onto a projection target (projection screen). In one embodiment, the basic video image is, for example, content written by the speaker on a blackboard or whiteboard. The enhanced video image includes the basic video image and the recognition result corresponding to the speaker's physical operations. The enhanced video image and control commands are transmitted to the electronic device 2. Electronic device 2 is used to provide video feeds to an external remote device 3 of a remote participant via a video conferencing platform. In one embodiment, electronic device 2 provides the received enhanced video feed as the video feed to the external remote device 3. In another embodiment, electronic device 2 performs control on a corresponding application according to received control commands to update the application display screen and provides the updated application display screen as the video feed to the external remote device 3. The application is, for example, word processing software, and this application is not limited thereto.
[0025] In one embodiment, the video processing terminal 1 is coupled to the electronic device 2 via a standard interface. The standard interface may be, for example, Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or wireless LAN communication technology, and this application is not limited thereto. In this way, the video processing terminal 1 can be recognized by the electronic device 2 as a camera, microphone, or HID device through the standard interface, achieving plug-and-play compatibility with existing electronic devices 2 without the need for additional driver installation. It can also be applied to various mainstream video conferencing platforms, significantly lowering the barrier to entry and enhancing ease of use.
[0026] Figure 2 This is a schematic diagram of another application environment of the video processing terminal according to an embodiment of this application. In this embodiment, the electronic device 2 is coupled to the projection device 4. The projection device 4 is used to receive image information from the electronic device 2, the image information including a basic image, and the projection device 4 projects the basic image onto the projection target (projection screen).
[0027] Please also refer to Figure 1 and Figure 3 , Figure 3 This is a schematic diagram of the system architecture of a video processing terminal according to an embodiment of this application. The video processing terminal 1 includes a processing module 110, a storage module 120, a communication module 130, an image capturing module 140, and a sound receiving module 150. The processing module 110 is coupled to the storage module 120, the communication module 130, the image capturing module 140, and the sound receiving module 150. The storage module 120 is used to store programs corresponding to the video processing system. The storage module 120 is, for example, a memory and / or a hard disk, but this application is not limited thereto. The processing module 110 is used to execute the programs of the video processing system to implement the corresponding video processing method. The processing module 110 is, for example, a processing circuit including a microprocessor, and this application is not limited thereto. The communication module 130 is coupled to the electronic device 2. In one embodiment, the communication module 130 is used to receive image information from the electronic device 2. In one embodiment, the communication module 130 is used to transmit control commands and / or enhanced image information to the electronic device 2. In one embodiment, the communication module 130 may be implemented using a standard interface circuit. The standard interface may be, for example, a Universal Serial Bus, a High Definition Multimedia Interface (HDM), or a Wireless Local Area Network (WLAN) communication technology, and this application is not limited thereto. The image-capturing module 140 is used to capture real-time local images including the speaker and transmit the real-time local images to the processing module 110. The image-capturing module 140 may be, for example, an image-capturing device such as a camera lens, and this application is not limited thereto. The audio-receiving module 150 is used to receive real-time audio information from the scene and transmit the real-time audio information to the processing module 110. The audio-receiving module 150 may be, for example, a microphone module, and this application is not limited thereto.
[0028] Figure 4 This is a schematic diagram of another system architecture of a video processing terminal according to an embodiment of this application. In this embodiment, the video processing terminal 1 further includes a projection module 160. The projection module 160 is coupled to the processing module 110. The projection module 160 is used to project image information from the electronic device 2 onto a projection target. The projection module 160 may be implemented by a projection circuit including a projection lens, and this application is not limited thereto.
[0029] Figure 5 This is a schematic diagram of the architecture of a video processing system according to an embodiment of this application. The video processing system includes an image processing module 191, a visual processing module 192, a voice processing module 193, a decision processing module 194, and an image processing module 195.
[0030] The image processing module 191 is used to extract image feature information from the image image information. The image feature information includes text information and object information in the basic image image. The image processing module 191 can detect text regions in the image image information and perform optical character recognition on the text regions to obtain text information. The image processing module 191 can perform object detection in the image image information, extract features from the detected objects (such as extracting features such as color, shape, and / or chart type), and generate object information corresponding to the objects.
[0031] In one embodiment, the image display information comes from electronic device 2. The image display information is the application display screen of electronic device 2. The image display information includes a basic image display. For example, the basic image display is the content of a presentation file displayed on electronic device 2. Text information is the text content in the displayed presentation file, and object information is the graphic content and / or object content in the displayed presentation file that can be selected, edited, moved, and formatted. In one embodiment, the image display information is a real-time local image, which includes content written by the speaker on a blackboard or whiteboard. The image processing module 191 identifies the image display area and the text information and / or object information within the image display area from the real-time local image. For example, the image processing module 191 identifies the area of the blackboard (image display area) and the text (text information) and / or graphics (object information) on the blackboard from the real-time local image.
[0032] In one embodiment, the image processing module 191 can also automatically correct text information. The image processing module 191 can automatically correct the format and / or typos of multiple text messages in the same image frame. For example, the image processing module 191 can identify multiple text messages and their formats (font size, font type, and / or paragraph settings, etc.) in the image frame and automatically change the format of the multiple text messages to a consistent format. The image processing module 191 can also automatically replace incorrect text content in the text information with correct text content based on a predefined database and / or a trained model.
[0033] The visual processing module 192 is used to extract visual feature information from the real-time local image. The visual processing module 192 is used to determine whether the speaker's image overlaps with the image frame area from the real-time local image. For example, the visual processing module 192 first individually identifies the speaker's image and the image frame area from the real-time local image, and confirms whether the speaker's image overlaps with the image frame area by determining whether the speaker's image falls within a specific area. When it is determined that the speaker's image overlaps with the image frame area, the visual processing module 192 generates overlap information. The overlap information may include coordinate information indicating the overlap between the speaker's image and the image frame area. Visual feature information includes overlap information.
[0034] The visual processing module 192 is used to determine the feature type of the real-time local image and confirm the corresponding indicator coordinates or gesture commands. In this embodiment, the visual feature information includes indicator coordinates and / or gesture commands. The visual processing module 192 is used to identify the speaker's hand area and / or indicator area from the real-time local image and determine the feature type of the real-time local image. When the visual processing module 192 determines that the real-time local image includes a hand area, it determines that the real-time local image is of the hand type. When the visual processing module 192 determines that the real-time local image includes an indicator area, it determines that the real-time local image is of the indicator type.
[0035] When the real-time local image is of the hand type, the vision processing module 192 further confirms the hand region. For example, the vision processing module 192 locates the hand region from the real-time local image based on feature extraction and bounding box regression to eliminate interference from non-hand regions. The vision processing module 192 also extracts hand key points. The vision processing module 192 outputs hand key points using a key point detection model based on skeletal topology. Hand key points include fingertips, knuckles, and the center of the palm. Thus, the vision processing module 192 can obtain fingertip coordinates from the hand key points, which are used as indicator coordinates for indicator points. In one embodiment, the fingertip coordinates of the index finger are used as indicator coordinates. In another embodiment, the indicator coordinates are determined by a preset order. For example, the preset order is index finger, middle finger, and thumb. When the fingertip coordinates of the index finger cannot be reliably detected, but the fingertip coordinates of the middle finger can be detected, the fingertip coordinates of the middle finger are used as indicator coordinates. In another embodiment, the fingertip coordinates detected by multiple consecutive real-time local images are used as the indicator coordinates. This allows the present application to determine the current position indicated by the speaker's hand.
[0036] The visual processing module 192 can further determine whether the hand area corresponds to a specific gesture. The visual processing module 192 determines whether the hand area corresponds to a specific gesture based on multiple consecutive real-time local images. For example, the visual processing module 192 confirms the movement trajectory based on the changes in the indicator coordinates of specific key points in multiple consecutive real-time local images, and determines the specific gesture based on the movement trajectory. Specific gestures include, for example, drawing a circle, moving away from the speaker in the real-time local image, or moving towards the speaker in the real-time local image. The visual processing module 192 then generates a gesture instruction based on predefined gesture instruction information corresponding to the specific gesture. The gesture instruction is used to indicate the corresponding control instruction for the specific gesture. For example, when the specific gesture is drawing a circle, the gesture instruction is a "circle selection" control instruction. As another example, when the specific gesture is moving away from the speaker, the gesture instruction is a "jump to the next page" control instruction.
[0037] When the real-time local image is an indicator, the vision processing module 192 further confirms the corresponding indicator coordinates of the indicator image. Indicators include, for example, a pointer, an object held by the speaker, or a pen tip. For instance, the vision processing module 192 can use an improved single-stage detection model to optimize the identification of indicators in the real-time local image and confirm the corresponding indicator coordinates of the indicator image, thereby improving the accuracy and performance of the identification. The vision processing module 192 can also further process the identified indicator image, such as removing overlapping detection boxes and combining indicator image features to filter out non-indicator images in the real-time local image, thereby improving the accuracy of the identification.
[0038] In one embodiment, the vision processing module 192 can also preprocess the real-time local image. For example, the vision processing module 192 can convert the real-time local image from BRG / RGB format to a model-compatible RGB format to eliminate device color deviation. The vision processing module 192 can suppress noise in the real-time local image by using Gaussian filtering combined with median filtering to remove salt-and-pepper noise and Gaussian noise. The vision processing module 192 can enhance the contrast of the real-time local image to improve the grayscale difference between the target area and the background. This improves the accuracy of determining the coordinates and / or gesture commands.
[0039] The speech processing module 193 is used to extract speech feature information from real-time speech information. The speech processing module 193 is used to identify the language type of the real-time speech information (e.g., Chinese, English, etc.), convert the real-time speech information into text content, and perform semantic analysis on the text content to generate speech feature information.
[0040] The decision processing module 194 determines the update content of the basic image frame based on image feature information, visual feature information, and voice feature information. The decision processing module 194 performs fusion judgment and annotation based on image feature information, visual feature information, and voice feature information to determine the update content. The update content is an instruction to change the basic image frame. In one embodiment, the update content includes indicator coordinates and their corresponding operation instructions. In one embodiment, the update content includes overlap information. In one embodiment, the update content includes gesture commands, and the update content is used to transmit to the electronic device 2.
[0041] Furthermore, the decision processing module 194 preprocesses the image feature information, visual feature information, and voice feature information. The decision processing module 194 performs basic validity verification and data organization on the image feature information, visual feature information, and voice feature information to prepare for subsequent fusion judgment.
[0042] In one embodiment, the decision processing module 194 verifies the accuracy of coordinate positioning and the completeness of content recognition of text and object information within the image feature information. The decision processing module 194 also marks the editable text area, object area, and unique identifier within the image information range to form an image feature coordinate content mapping table.
[0043] In one embodiment, the decision processing module 194 verifies the validity of the primitives of the indicator coordinates, the gesture of the gesture command, the integrity of the indicator trajectory, and the accuracy of the coordinate range of the overlapping information. It removes blurry and / or meaningless visual features (such as fingertip coordinates that fail to be detected or invalid gestures without trajectories) and adds timestamps to the valid features.
[0044] In one embodiment, the decision processing module 194 verifies the completeness of the semantic analysis results of the speech feature information, extracts core keywords / instructions (such as "highlight", "select", "next page", "peak value"), marks the intent type of the speech semantics (such as screen operation, page control, content emphasis), and adds a timestamp synchronized with the real-time speech.
[0045] By preprocessing image feature information, visual feature information, and voice feature information, three types of effective feature information can be retained, invalid data can be eliminated, and the timestamp format and spatial coordinate reference (all mapped to the primitive coordinate system of the basic image) can be unified, so as to effectively improve the accuracy of the updated content.
[0046] Furthermore, the decision processing module 194 performs spatiotemporal alignment on the image feature information, visual feature information, and voice feature information. The decision processing module 194 performs dual alignment of the image feature information, visual feature information, and voice feature information through time synchronization and spatial mapping to ensure consistency in the fusion judgment.
[0047] In one embodiment, the decision processing module 194 performs time alignment on image feature information, visual feature information, and voice feature information. Based on the acquisition timestamp of real-time local image and / or real-time voice information, the decision processing module 194 synchronizes the parsing time of image feature information, the recognition time of visual feature information, and the conversion time of voice feature information at the millisecond level, retaining only the three types of feature information within the same time window for fusion judgment, thus avoiding invalid linkage across time periods.
[0048] In one embodiment, the decision processing module 194 performs spatial mapping on image feature information, visual feature information, and speech feature information. It precisely maps the indicator coordinates / overlapping information coordinates in the visual feature information to the primitive coordinate system of the image feature coordinate content mapping table. It performs preliminary semantic matching between speech keywords in the speech feature information and text / object information in the image feature information, and marks the image feature candidate regions corresponding to the speech keywords.
[0049] By aligning image, visual, and voice features in a spatiotemporal dimension, a multi-source feature fusion dataset with "the same time window and the same spatial coordinate reference" is formed. This enables three-dimensional linkage mapping of image, visual, and voice features, effectively improving the accuracy of updated content.
[0050] In one embodiment, the decision processing module 194 determines the primary feature type based on the real-time local image. The decision processing module 194 determines the corresponding primary feature type of the real-time local image based on the real-time local image. For example, when the real-time local image shows a speaker pointing at the screen with their hand and / or an indicator, the decision processing module 194 determines that the primary feature type of the real-time local image is a visual feature based on the indicator coordinates in the visual feature information. For example, when the real-time local image shows a speaker emphasizing the content of the screen with voice, the decision processing module 194 determines that the primary feature type of the real-time local image is a voice feature based on the semantic keywords in the voice feature information. For example, when the real-time local image shows a speaker obscuring the screen, the decision processing module 194 determines that the primary feature type of the real-time local image is both a screen feature and a visual feature based on the overlap information in the visual feature information. For example, when the real-time local image shows a speaker combining gestures and voice, the decision processing module 194 determines that the primary feature type of the real-time local image is both a visual feature and a voice feature based on the gesture command information in the visual feature information and the semantic keywords in the voice feature information. For example, when there is no speaker or the speaker does not perform any physical operation in the real-time local image, the decision processing module 194 determines that the main feature type of the real-time local image is the image feature.
[0051] In one embodiment, the decision processing module 194 assigns weights to image feature information, visual feature information, and voice feature information based on the primary feature type. The decision processing module 194 assigns weights to image feature information, visual feature information, and voice feature information according to the determined primary feature type using a predefined weight allocation table. Table 1 shows an embodiment of the weight allocation table in this application. For example, when the decision processing module 194 determines that the primary feature type is visual feature, the decision processing module 194 sets the weight of image feature information to 20%, the weight of visual feature information to 70%, and the weight of voice feature information to 10% according to the weight allocation table.
[0052] Table 1
[0053] Image feature information weight Visual feature information weights Speech feature information weights Main feature types 20% 70% 10% Visual features 20% 10% 70% speech features 50% 50% 0% Image and visual features 20% 40% 40% Visual and speech features 100% 0% 0% Image features
[0054] In one embodiment, the decision processing module 194 determines the update content based on the weights of the main feature type and the image feature information, visual feature information, and voice feature information. The decision processing module 194 performs scene-specific fusion judgment on three types of update content: overlapping information, gesture commands, and coordinate operation instructions. Combined with the linkage verification of main features and auxiliary features, the module completes the initial annotation of the update content.
[0055] Furthermore, the decision processing module 194 determines the initial update content based on the main feature type, screen feature information, visual feature information, and voice feature information.
[0056] In the embodiment where the speaker and the image frame overlap, the decision processing module 194 determines that the main feature is a visual feature based on the main feature type. Therefore, it prioritizes the overlap coordinate information in the visual feature information (weight ≥ 50%), while combining the image frame range coordinates in the image feature information as auxiliary features to confirm the specific coordinates of the occluded area, the occluded area, and the corresponding image text / object information. Simultaneously, the decision processing module 194 initially labels the primitive coordinate range of the overlapping area, the unique identifier of the occluded image text / object, and the occlusion level to form standardized overlap information (initial update content).
[0057] In embodiments where the speaker operates using gestures and / or pointers, the decision processing module 194 uses the gesture / pointer trajectory features from visual feature information as the primary feature (weight ≥ 40%), combined with the semantic intent of voice feature information as an auxiliary feature (weight ≥ 40%), and screen feature information as an auxiliary feature for scene adaptability verification (weight 20%). Simultaneously, the decision processing module 194 initially labels the gesture command type (e.g., page turning, selection, scrolling), command execution trigger conditions, and the corresponding electronic device / screen control operation, forming standardized gesture commands (initial update content). In this embodiment, the standardized gesture commands can be directly transmitted to the electronic device 2. For example, the decision processing module 194 extracts the gesture / pointer trajectory type (e.g., drawing a circle, moving away from the speaker, moving closer to the speaker, swiping up and down) based on visual feature information, matches it with a predefined set of gesture command candidates (e.g., selection, next page, previous page, scrolling), and extracts semantic command keywords (e.g., "next page," "select," "zoom in") from voice feature information, performing semantic matching with the gesture command candidate set. The decision processing module 194 performs image feature verification to confirm whether the current basic image supports the instruction (e.g., whether there is a next page of content, whether it can be zoomed in). The decision processing module 194 only confirms the gesture instruction if the trajectory of the gesture / pointer matches the speech semantics with a degree greater than or equal to a preset threshold; otherwise, it discards the candidate instruction.
[0058] In embodiments where the speaker uses gestures and / or pointers to indicate the screen, the decision processing module 194 uses the pointer coordinates in the visual feature information as the primary feature (weight ≥ 40%), combines the semantic keywords of the voice feature information to determine the operation type (weight ≥ 40%), and uses the screen feature information to perform coordinate content matching (weight 20%). Simultaneously, the decision processing module 194 initially labels the precise primitive coordinates of the pointer point, the unique identifier of the corresponding screen text / object, and the matched screen operation instruction type (e.g., highlight, wrap, label, zoom), forming a combined update of "pointer coordinates + operation instruction". For example, the decision processing module 194 maps the pointer coordinates of the visual feature information to the screen feature coordinate content mapping table of the screen feature information to accurately locate the screen text / object information corresponding to the pointer coordinates. The decision processing module 194 uses the voice feature information to extract the screen text / object information corresponding to the content emphasis / operation (e.g., "highlight", "label", "emphasize", "zoom") to determine the corresponding screen operation instruction type. The decision processing module 194 also uses the screen feature information to verify whether the text / object supports the operation type (e.g., whether it can be highlighted, whether it can be labeled).
[0059] Furthermore, the decision processing module 194 verifies the initial update content. In this step, the decision processing module 194 performs cross-feature conflict verification on the annotated initial update content and resolves conflicts according to preset rules to ensure the accuracy and uniqueness of the updated content, thereby avoiding situations such as inconsistencies between indicated coordinates and voice semantics, and conflicts between gesture commands and screen content.
[0060] In one embodiment, the decision processing module 194 is used to identify typical conflicts. For example, the coordinates of the visual feature information are inconsistent with the content of the image pointed to by the voice semantics, the gesture command conflicts with the operable attributes of the basic image, and the occlusion area marked by overlapping information does not match the actual image features.
[0061] In one embodiment, the decision processing module 194 is used to resolve conflicts. For example, in a teaching scenario, if the deviation of the indicated coordinates is small and the voice semantics are clear, the indicated coordinates are corrected based on the voice semantics. For example, in a meeting scenario, if the indicated coordinates are accurate but the voice semantics are ambiguous, the indicated coordinates are used as the standard to match the preset operation of the screen features (such as highlighting). For example, if a gesture command conflicts with the operable attributes of the screen (such as determining a page-turning command when there is no next page), the gesture command is directly rejected and marked as invalid.
[0062] In one embodiment, the decision processing module 194 performs a secondary validity check. The decision processing module 194 re-verifies the compatibility of the conflict-resolved update content with the basic image frame, eliminating invalid / unexecutable update content. This yields a conflict-free and executable final set of candidate update content.
[0063] Furthermore, the decision processing module 194 generates and outputs the final updated content. The decision processing module 194 converts the candidate set of updated content after resolving conflicts into a standardized format recognizable by the image processing module 195 / electronic device 2, and classifies and outputs it according to the type of updated content.
[0064] In one embodiment, the decision processing module 194 uniformly converts the coordinates, instructions, and operation types of all updated content into the primitive coordinate system and instruction protocol of the screen processing module 195, and the control instructions of the electronic device 2 are uniformly converted into the standard communication protocol.
[0065] In one embodiment, the decision processing module 194 transmits the updated content of the overlap information, indicator coordinates, and operation instructions to the screen processing module 195 as the basis for generating the enhanced image screen. The decision processing module 194 transmits the updated content of the gesture instructions to the electronic device 2 to control the application display screen of the electronic device 2, or synchronously transmits it to the screen processing module 195 for screen linkage.
[0066] In one embodiment, the decision processing module 194 records the multi-source feature information, weight allocation, conflict resolution process, and final update content for each decision, providing data support for subsequent feature recognition model optimization.
[0067] Through the above, the decision processing module 194 finally outputs unique, executable, and standardized update content. All update content achieves linkage matching of three types of features: visual, voice, and image, ensuring that the update content is highly consistent with the speaker's actual operation / narration intention, providing accurate basis for the subsequent image processing module 195 to generate enhanced image images and for the electronic device 2 to perform control operations.
[0068] The image processing module 195 generates an enhanced image based on the updated content and the basic image. The enhanced image is then transmitted to the electronic device 2.
[0069] In one embodiment, the image processing module 195 generates a basic image frame using text information and object information. In embodiments where the image frame information is a real-time local image, including content written by the speaker on a blackboard or whiteboard, the image processing module 195 generates the basic image frame based on the text information and / or object information in the image frame information. For example, the image processing module 195 converts the coordinate information of the text information and / or object information within the image frame area into primitive coordinates corresponding to a preset image (e.g., a blank image), and overlays the text information and / or object information onto the preset image based on the primitive coordinates to generate the basic image frame.
[0070] In one embodiment, the image processing module 195 further generates a basic image frame based on the overlap information. The image processing module 195 can generate the current basic image frame based on the overlap information and historical basic image frames, using the content (text information and / or object information) of the historical basic image frames corresponding to the overlap information. In this way, obscured content can be compensated for by historical basic image frames, generating an unobscured basic image frame, allowing the remote participant to receive all content of the basic image frame completely.
[0071] The image processing module 195 is used to confirm the primitive coordinates of the indicator point according to the updated content, adjust the parameter configuration of the indicator point, and generate an enhanced image screen including the indicator point based on the primitive coordinates and parameter configuration. In this embodiment, the image processing module 195 confirms the indicator coordinates of the indicator point according to the updated content and converts the indicator coordinates into primitive coordinates corresponding to the basic image screen. The image processing module 195 also adjusts the parameter configuration of the indicator point. The parameter configuration includes, for example, the radius, color, and other display settings of the indicator point. Based on the corresponding primitive coordinates and parameter configuration, the image processing module 195 overlays the indicator point onto the basic image screen to generate an enhanced image screen including the indicator point. Thus, remote participants can clearly see the current position indicated by the speaker on the basic image screen through the enhanced image screen.
[0072] In one embodiment, the image processing module 195 overlays indicator points onto a series of basic image frames based on the speaker's pointing trajectory, thereby generating an indicator line formed by the continuous indicator points. This enhances the visual effect of dynamic following.
[0073] In one embodiment, the screen processing module 195 is used to enable historical coordinate caching and generate transition coordinates through linear interpolation, which are then used as primitive coordinates. When the screen processing module 195 does not obtain the indicator coordinates, it can generate transition coordinates through historical indicator coordinates and use the transition coordinates as the primitive coordinates of the indicator point to avoid the indicator point from disappearing or changing.
[0074] In one embodiment, the image processing module 195 performs trajectory smoothing on the primitive coordinates. The image processing module 195 can use a Kalman filter algorithm to perform trajectory smoothing on the primitive coordinates, ensuring the trajectory tracking of the primitive coordinates. The image processing module 195 dynamically adjusts the noise weight based on the confidence level. In one embodiment, when the movement speed of the primitive coordinates in a continuous basic image frame is less than a preset threshold (e.g., 5 primitives / frame), the image processing module 195 performs trajectory smoothing on the primitive coordinates using a moving average filter algorithm, thereby reducing static jitter.
[0075] In one embodiment, the image processing module 195 determines the location information and operation type based on the updated content, determines the corresponding text information and / or object information based on the location information, and adjusts the visual effects of the text information and / or object information according to the operation type to generate an enhanced image. The location information may be, for example, gesture commands and / or voice feature information. The operation type may be, for example, highlighting, text wrapping, object wrapping, etc. After determining the location information and operation type based on the updated content, the image processing module 195 compares the text information and object information in the basic image with the gesture commands and / or voice feature information, determines the corresponding text information and / or object information based on the comparison result, and changes the visual effects of the text information and / or object information according to the operation type. In one embodiment, the accuracy of the comparison result can determine the corresponding text information and / or object information. In another embodiment, the relevance of the comparison result can determine the corresponding text information and / or object information. For example, the image processing module 195 determines the gesture command as "selection" and the operation type as highlighting based on the updated content. The image processing module 195 then confirms the selected area based on the corresponding coordinates of the gesture command, adjusts the display effect of the text information within the area to highlight, and generates an enhanced image. For example, the image processing module 195 determines the speech feature information as "peak value" and the operation type as highlighting based on the updated content. The image processing module 195 then adjusts the display effect of the peak area of the waveform object information corresponding to the speech feature information to highlight, thereby generating an enhanced image. In this way, remote participants can clearly understand the content currently emphasized by the speaker on the basic image through the enhanced image.
[0076] Figure 6 This is a flowchart illustrating a video processing method according to an embodiment of this application, which can be implemented by the aforementioned video processing terminal 1 and video processing system. The video processing method includes steps S1 to S4.
[0077] Step S1: Obtain video image information, real-time local video, and real-time audio information. In this step, the image capturing module 140 of the video processing terminal 1 is used to capture real-time local video including the speaker, and the audio receiving module 150 of the video processing terminal 1 is used to receive real-time audio information from the scene. In one embodiment, the communication module 130 of the video processing terminal 1 is used to receive video image information from the electronic device 2. In one embodiment, the real-time local video is used as video image information.
[0078] Step S2: Obtain image feature information, visual feature information, and voice feature information. In this step, the image processing module 191 extracts image feature information from the image image information, which includes text information and object information in the basic image image. The visual processing module 192 extracts visual feature information from the real-time local image. The visual feature information includes the coordinates of the indicator points, gesture commands, and / or overlap information. Further, the visual processing module 192 determines the feature type of the real-time local image and confirms the indicator coordinates or gesture commands corresponding to the feature type. The voice processing module 193 extracts voice feature information from the real-time voice information.
[0079] Step S3: Determine the video feed update content based on image feature information, visual feature information, and voice feature information. In this step, the decision processing module 194 performs fusion judgment and annotation based on image feature information, visual feature information, and voice feature information to determine the update content. The update content is an instruction to change the basic image feed. In one embodiment, when the update content is a gesture command, the update content is transmitted to the electronic device 2. Thereby, the electronic device 2 can control the corresponding application according to the received control command to update the application display screen and provide the updated application display screen as the video feed to the remote participant.
[0080] Step S4: Generate an enhanced image based on the updated content and the basic image. In this step, the image processing module 195 confirms the changes to be made to the basic image based on the updated content and generates a corresponding enhanced image based on the basic image. The enhanced image is then transmitted to the electronic device 2, which provides the received enhanced image as a video feed to the remote participant. In one embodiment, the image processing module 195 also generates the basic image based on text and object information. In another embodiment, the image processing module 195 further generates the basic image based on overlap information.
[0081] In one embodiment, step S3 includes steps S31 to S34, such as... Figure 7 As shown.
[0082] Step S31: Preprocess the image feature information, visual feature information, and voice feature information. In this step, the decision processing module 194 performs basic validity verification and data organization on the image feature information, visual feature information, and voice feature information to prepare for subsequent fusion judgment.
[0083] Step S32: Perform spatiotemporal alignment of image feature information, visual feature information, and speech feature information. In this step, the decision processing module 194 performs dual alignment of image feature information, visual feature information, and speech feature information through temporal synchronization and spatial mapping to ensure consistency in the fusion judgment.
[0084] Step S33: Determine the main feature type based on the real-time local image. In this step, the decision processing module 194 determines the corresponding main feature type of the real-time local image based on the real-time local image.
[0085] Step S34: Based on the main feature types, assign weights to the image feature information, visual feature information, and speech feature information. In this step, the decision processing module 194 assigns weights to the image feature information, visual feature information, and speech feature information according to the determined main feature types using a predefined weight allocation table.
[0086] Step S35: Determine the update content based on the weights of the main feature type and image feature information, visual feature information, and voice feature information. In this step, the decision processing module 194 performs scene-specific fusion judgment on three types of update content: overlapping information, gesture commands, and coordinate operation instructions. Combined with the linkage verification of main features and auxiliary features, the initial annotation of the update content is completed.
[0087] In one embodiment, step S35 includes steps S351 to S353, such as... Figure 8 As shown.
[0088] Step S351: Determine the initial update content based on the main feature type, image feature information, visual feature information, and voice feature information.
[0089] Step S352: Verify the initial update content. In this step, the decision processing module 194 performs cross-feature conflict verification on the initially updated content and resolves conflicts according to preset rules to ensure the accuracy and uniqueness of the updated content, thereby avoiding situations such as inconsistencies between indicated coordinates and voice semantics, and conflicts between gesture commands and screen content.
[0090] In one embodiment, the decision processing module 194 is used to identify typical conflicts, resolve conflicts, and perform a secondary validity check. This yields a conflict-free, executable final update content candidate set.
[0091] Step S353: Generate and output the final updated content. In this step, the decision processing module 194 converts the conflict-resolved candidate set of updated content into a standardized format recognizable by the screen processing module 195 / electronic device 2, and classifies and outputs it according to the type of updated content.
[0092] In one embodiment, the decision processing module 194 uniformly converts the coordinates, instructions, and operation types of all updated content into the primitive coordinate system and instruction protocol of the screen processing module 195, and the control instructions of the electronic device 2 are uniformly converted into the standard communication protocol.
[0093] In one embodiment, the decision processing module 194 transmits overlapping information, indicator coordinates, and operation instruction-type update content to the screen processing module 195 as the basis for generating enhanced image screens. The decision processing module 194 transmits gesture command-type update content to the electronic device 2 to control the application display screen of the electronic device 2, or synchronously transmits it to the screen processing module 195 for screen linkage.
[0094] In one embodiment, the decision processing module 194 records the multi-source feature information, weight allocation, conflict resolution process, and final update content for each decision, providing data support for subsequent feature recognition model optimization.
[0095] In one embodiment, step S4 includes steps S411 to S415, such as... Figure 9 As shown.
[0096] Step S411: Confirm the element coordinates of the indicator point based on the updated content. In this step, the image processing module 195 obtains the indicator coordinates of the indicator point based on the updated content and converts the indicator coordinates into element coordinates corresponding to the basic image frame.
[0097] Step S413: Adjust the parameter configuration of the indicator point. In this step, the screen processing module 195 adjusts the parameter configuration of the indicator point. The parameter configuration includes, for example, the radius and color of the indicator point, as well as other display settings for the indicator point.
[0098] Step S415: Generate an enhanced image image including indicator points based on the primitive coordinates and parameter configuration. In this step, the image processing module 195 overlays the indicator points onto the basic image image according to the corresponding primitive coordinates and parameter configuration to generate an enhanced image image including the indicator points.
[0099] In one embodiment, step S4 further includes step S414, such as Figure 10 As shown.
[0100] Step S414: Abnormal Indicator Point Handling. In this step, the image processing module 195 enables historical coordinate caching and generates transition coordinates through linear interpolation. These transition coordinates serve as the primitive coordinates. Historical coordinates may include, for example, historical indicator coordinates.
[0101] In one embodiment, step S4 includes steps S421 to S415, such as... Figure 11 As shown.
[0102] Step S421: Determine location information and operation type based on the updated content. In this step, the image processing module 195 determines the location information and operation type based on the received updated content. The location information may be, for example, gesture commands and / or voice feature information. The operation type may be, for example, highlighting, text wrapping, object wrapping, etc.
[0103] Step S422: Determine the corresponding text information and / or object information based on the positioning information. In this step, the image processing module 195 compares the text information and object information in the basic image frame with gesture commands and / or voice feature information, and determines the corresponding text information and / or object information based on the comparison results. In one embodiment, the accuracy of the comparison results can determine the corresponding text information and / or object information. In another embodiment, the relevance of the comparison results can determine the corresponding text information and / or object information.
[0104] Step S423: Update the basic image screen according to the operation type to generate an enhanced image screen. After determining the corresponding text information and / or object information, the image processing module 195 changes the visual effects of the text information and / or object information according to the operation type to generate an enhanced image screen.
[0105] In the embodiments of this application, Figures 7 to 11 The execution order of the steps can be adjusted according to the requirements, and can be executed simultaneously or individually, and this application is not limited to this.
[0106] In summary, the embodiments of this application, through a decision processing module, jointly determine the update content of the basic image frame based on screen feature information, visual feature information, and voice feature information. Based on the updated content and the basic image frame, an enhanced image frame is automatically generated according to the speaker's physical operation and / or narration content. This allows remote participants to clearly understand the speaker's key points through the enhanced image frame, improving communication efficiency and information transmission accuracy, effectively increasing the interactivity between the speaker and remote participants, and enhancing the clarity of the speaker's instructions.
[0107] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0108] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms fall within the scope of protection of this application.
Claims
1. A video image processing terminal, characterized in that, include: The image processing module extracts image feature information from the image image information, and the image feature information includes text information and object information; The visual processing module extracts visual feature information from real-time local images; The speech processing module extracts speech feature information from real-time speech information; The decision processing module determines the update content of the basic image frame based on the image feature information, the visual feature information, and the voice feature information; and The image processing module generates an enhanced image based on the updated content and the basic image, and transmits the enhanced image as a video feed to an external remote device.
2. The video processing terminal according to claim 1, characterized in that, The image information is the application display screen of the electronic device, and the image information includes the basic image screen.
3. The video processing terminal according to claim 1, characterized in that, The image information is the real-time local image. The image processing module receives the real-time local image and identifies the image area and the text information and object information within the image area from the real-time local image.
4. The video processing terminal according to claim 1, characterized in that, The image processing module corrects multiple text messages in the image information.
5. The video processing terminal according to claim 1, characterized in that, The visual processing module is used to determine the feature type of the real-time local image and confirm the indicator coordinates or gesture commands corresponding to the feature type. The visual feature information includes the indicator coordinates and / or the gesture commands.
6. The video processing terminal according to claim 1, characterized in that, The speech processing module is used to convert the real-time speech information into text content and perform semantic analysis on the text content to generate the speech feature information.
7. The video processing terminal according to claim 1, characterized in that, The image processing module is used to confirm the primitive coordinates of the indicator point according to the updated content, adjust the parameter configuration of the indicator point, and generate the enhanced image image including the indicator point based on the primitive coordinates and the parameter configuration.
8. The video processing terminal according to claim 7, characterized in that, The image processing module is used to enable historical coordinate caching and generate transition coordinates through linear interpolation, and the transition coordinates are used as the coordinates of the graphic element.
9. The video processing terminal according to claim 1, characterized in that, The image processing module is used to determine the location information and operation type according to the updated content, and to determine the corresponding text information and / or object information according to the location information, and to adjust the visual effect of the text information and / or object information according to the operation type, so as to generate the enhanced image.
10. The video processing terminal according to claim 1, characterized in that, The image processing module generates the basic image based on the text information and the object information.
11. The video processing terminal according to claim 1, characterized in that, The video processing terminal is coupled to electronic devices via a standard interface.
12. A video image processing method, characterized in that, include: Image feature information is extracted from image image information, including text information and object information; visual feature information is extracted from real-time local images; and voice feature information is extracted from real-time voice information. The update content of the basic image frame is determined based on the image feature information, the visual feature information, and the voice feature information; as well as An enhanced image is generated based on the updated content and the basic image, and the enhanced image is used as a video feed to be transmitted to an external remote device.
13. The video image processing method according to claim 12, characterized in that, The image information is the software display screen of the electronic device, and the image information includes the basic image screen.
14. The video image processing method according to claim 12, characterized in that, When the image information is the real-time local image, the image range and the text information and object information within the image range are identified from the real-time local image.
15. The video image processing method according to claim 12, characterized in that, The basic image is generated based on the text information and the object information.
16. The video image processing method according to claim 12, characterized in that, The feature type of the real-time local image is determined, and the corresponding indicator coordinates or gesture commands are confirmed. The visual feature information includes the indicator coordinates and / or the gesture commands.
17. The video image processing method according to claim 12, characterized in that, The real-time voice information is converted into text content, and semantic analysis is performed on the text content to generate the voice feature information.
18. The video image processing method according to claim 12, characterized in that, The video processing method includes: Confirm the element coordinates of the indicated point based on the updated content; Adjust the parameter configuration of the indicator point; and The enhanced imagery, including the indicator points, is generated based on the primitive coordinates and the parameter configuration.
19. The video image processing method according to claim 18, characterized in that, Enable historical coordinate caching, generate transition coordinates through linear interpolation, and use the transition coordinates as the coordinates of the primitive.
20. The video image processing method according to claim 12, characterized in that, The video processing method includes: The location information and operation type are determined based on the updated content; Determine the corresponding text information and / or object information based on the location information; and The basic image frame is updated according to the operation type to generate the enhanced image frame.
21. A video image processing terminal, characterized in that, It includes: Memory, which stores the program of the video processing system; An image acquisition device used to obtain real-time local images; A microprocessor, coupled to the memory and the image capturing device, is used to execute the program of the video image processing system to implement a corresponding video image processing method, the video image processing method including: Image feature information is extracted from image image information, including text information and object information; visual feature information is extracted from the real-time local image; and voice feature information is extracted from the real-time voice information. The update content of the basic image frame is determined based on the image feature information, the visual feature information, and the voice feature information; as well as An enhanced image is generated based on the updated content and the basic image, and the enhanced image is used as a video feed to be transmitted to an external remote device.
22. The video processing terminal according to claim 21, characterized in that, The image information is the software display screen of the electronic device, and the image information includes the basic image screen.
23. The video processing terminal according to claim 21, characterized in that, The video processing method includes: when the video information is the real-time local video, identifying the video range and the text information and object information within the video range from the real-time local video.
24. The video processing terminal according to claim 21, characterized in that, The video processing method includes generating the basic image based on the text information and the object information.
25. The video processing terminal according to claim 21, characterized in that, The video processing method includes: determining the feature type of the real-time local image and confirming the indicator coordinates or gesture commands corresponding to the feature type, wherein the visual feature information includes the indicator coordinates and / or the gesture commands.
26. The video processing terminal according to claim 21, characterized in that, The video processing method includes: converting the real-time audio information into text content, and performing semantic analysis on the text content to generate the audio feature information.
27. The video processing terminal according to claim 21, characterized in that, The video processing method includes: Confirm the element coordinates of the indicated point based on the updated content; Adjust the parameter configuration of the indicator point; and The enhanced imagery, including the indicator points, is generated based on the primitive coordinates and the parameter configuration.
28. The video processing terminal according to claim 27, characterized in that, The video processing method includes: enabling historical coordinate caching, generating transition coordinates through linear interpolation, and using the transition coordinates as the coordinates of the graphic element.
29. The video processing terminal according to claim 21, characterized in that, The video processing method includes: The location information and operation type are determined based on the updated content; Determine the corresponding text information and / or object information based on the location information; and The basic image frame is updated according to the operation type to generate the enhanced image frame.