A conference camera target detection method and system based on image recognition
Patent Information
- Application Number
- CN202610848260.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-15
Smart Images

Figure CN122765152A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular to a method and system for target detection in a conference camera based on image recognition. Background Technology
[0002] In modern meeting environments, intelligent conference camera systems are widely used to automatically detect and accurately track key targets in a meeting, such as speakers or presentation devices, through image recognition technology, thereby enhancing the experience for remote participants. However, existing intelligent conference camera systems face significant technical challenges when handling presentations of physical objects during meetings.
[0003] Specifically, in some professional conferences, the focus of discussion may not be limited to human-screen interaction but may also involve the demonstration and discussion of physical objects. For example, in a product design review meeting, a designer might place a newly designed prototype, a precision circuit board, or a complex mechanical model in the center of the conference table. At this point, the focus of the meeting shifts from the virtual content on the screen to this specific physical object. The designer might pick up the prototype, demonstrate it from multiple angles, point out its specific structural, material, or functional details, and may even conduct a live disassembly or assembly demonstration. For existing smart camera systems, their image recognition programs typically focus on faces, body movements, and large, static display devices. When faced with a relatively small, irregularly shaped physical object that may be frequently manipulated, the system's ability to recognize and continuously track it becomes insufficient. It may fail to accurately identify the physical object as a "key target," or even if it does, it may struggle to track it as steadily as it does with a face. The system may mistakenly focus on the presenter's hands or just a certain area of the table, failing to clearly capture the key details of the physical object being discussed. This prevents remote participants from seeing the specific features of the object and severely impacts their participation in the product discussion.
[0004] Furthermore, the problem becomes even more pronounced if the physical object being demonstrated is highly dynamic, intricate, or easily occluded. For example, an engineer might be demonstrating a tiny sensor module that is extremely small and requires frequent handling to show different sides, potentially even being partially obscured by the engineer's fingers. In such cases, the object's small size, high speed, varied postures, and susceptibility to occlusion pose a significant challenge to the real-time target detection and tracking algorithms of the image recognition system. The system may frequently "lose" the target within a short period, causing unnatural camera shakes or sudden switching as it struggles to recapture the target. Worse still, the system may be unable to distinguish whether to focus on the engineer's hand movements or on the miniature sensor itself. This unstable tracking and inaccurate focusing make it difficult for remote participants to maintain focus on the key points of the presentation, and the discontinuity and incoherence of the image significantly diminish the effectiveness of the demonstration. Summary of the Invention
[0005] This application provides a target detection method and system for conference cameras based on image recognition. It aims to solve the technical problems of existing intelligent conference camera systems in handling the demonstration of physical objects in a conference. These problems include the difficulty in accurately identifying and stably tracking physical objects that are relatively small in size, irregular in shape, and may be frequently manipulated, as well as the insufficient real-time target detection and tracking capabilities of the system when the physical object has highly dynamic, detailed, or easily occluded characteristics. As a result, remote participants cannot see the details of the object clearly, which affects their participation, and the image appears discontinuous and incoherent.
[0006] The technical solution of this application is as follows:
[0007] In a first aspect, this application discloses a target detection method for a conference camera based on image recognition, comprising: acquiring a behavioral image of a speaker; the behavioral image includes a gesture direction and a gaze direction; the gesture direction is the direction pointed by the speaker's index finger; determining whether the speaker intends to demonstrate a physical object based on the behavioral image, and when the speaker intends to demonstrate a physical object, controlling the conference camera to point at the area where the physical object is located; determining the object attribute information of the physical object based on the original image information of the physical object, adjusting the original image attribute information of the conference camera to obtain adjusted image attribute information; during the process of the speaker operating the physical object, displaying the physical object in the display screen based on the adjusted image attribute information according to the relative motion relationship between the speaker's hand and the physical object; during the display of the physical object, controlling the conference camera to adjust the image so that the physical object remains in the center of the screen.
[0008] Furthermore, based on the behavioral image, it is determined whether the speaker intends to demonstrate a physical object. When the speaker intends to demonstrate a physical object, the conference camera is controlled to be pointed at the area where the physical object is located. This includes: determining whether the direction of the gesture and the direction of the gaze intersect, and whether the intersection point is within the field of view of the conference camera; when the direction of the gesture and the direction of the gaze intersect, and the intersection point is within the field of view of the conference camera, an image recognition algorithm is used to determine whether a physical object is identified at the intersection point; when a physical object is identified at the intersection point, it is determined that the speaker intends to demonstrate a physical object, the area where the intersection point exists is taken as the area where the physical object is located, and the conference camera is controlled to be pointed at the area where the physical object is located; when the direction of the gesture and the direction of the gaze do not intersect, or when the direction of the gesture and the direction of the gaze intersect, and the intersection point is not within the field of view of the conference camera, or when the direction of the gesture and the direction of the gaze intersect, and the intersection point is within the field of view of the conference camera, and no physical object is identified at the intersection point, it is determined that the speaker does not intend to demonstrate a physical object.
[0009] More specifically, in some implementation schemes, when a physical object is identified at an intersection, it is determined that the speaker intends to demonstrate the physical object. The area where the intersection exists is taken as the area where the physical object is located, and the conference camera is controlled to be pointed at the area where the physical object is located. This includes: acquiring the speaker's voice information; converting the voice information into text to obtain text information, and determining whether there are preset words in the text information; when preset words are found in the text information, it is determined that the speaker intends to demonstrate the physical object, the area where the intersection exists is taken as the area where the physical object is located, and the conference camera is controlled to be pointed at the area where the physical object is located.
[0010] Building upon the above, this application further proposes that the object attribute information includes the color and movement speed of the physical object, and the original image attribute information includes the original exposure value and the original frame rate. The object attribute information of the physical object is determined based on its original image information, and the original image attribute information of the conference camera is adjusted to obtain the adjusted image attribute information. This includes: determining the color of the physical object based on its original image information, and determining the hue value of the physical object based on its color; using the product of the hue value and a preset color adjustment coefficient as the exposure compensation gain; adjusting the original exposure value of the conference camera based on the exposure compensation gain to obtain the adjusted exposure value; and determining the movement speed of the physical object based on its original image information, and adjusting the original frame rate of the conference camera based on the movement speed of the physical object to obtain the adjusted frame rate.
[0011] Preferably, adjusting the original exposure value of the conference camera based on the exposure compensation gain to obtain the adjusted exposure value includes: acquiring the light brightness of the conference room detected by the light sensor and a first correspondence; the first correspondence includes a one-to-one correspondence between multiple light brightness ranges and multiple light adjustment coefficients; using the light adjustment coefficient corresponding to the light brightness range in which the light brightness of the conference room is located as the target light adjustment coefficient; using the product of the exposure compensation gain and the target light adjustment coefficient as the target exposure compensation gain; and using the sum of the original exposure value of the conference camera and the target exposure compensation gain as the adjusted exposure value of the conference camera.
[0012] In one implementation, adjusting the frame rate of the original video feed of a conference camera based on the movement speed of a physical object includes: obtaining a second correspondence and an importance index of the current meeting; the second correspondence includes a one-to-one correspondence between multiple movement speed ranges and multiple frame rate adjustment coefficients; using the multiple frame rate adjustment coefficients corresponding to the movement speed range of the physical object in the second correspondence as initial frame rate adjustment coefficients; using the weighted sum of the initial frame rate adjustment coefficients and the importance index of the current meeting as a target frame rate adjustment coefficient; and using the product of the original video frame rate of the conference camera and the target frame rate adjustment coefficient as the adjusted video frame rate of the conference camera.
[0013] As a technological improvement, during the process of the speaker manipulating the physical object, the physical object is displayed on the screen based on the adjusted image attribute information according to the relative motion relationship between the speaker's hand and the physical object. This includes: determining whether the speaker's hand is obscuring the physical object; when the speaker's hand obscures the physical object, determining the obscured part of the physical object in the next frame image based on the relative motion relationship between the speaker's hand and the physical object; and filling the obscured part of the physical object in the next frame image based on the pre-captured image of the physical object, so as to display the physical object on the screen based on the adjusted image attribute information.
[0014] As a further improvement, the relative motion relationship includes motion speed and motion direction. Based on the relative motion relationship between the speaker's hand and the physical object, the part of the physical object that is occluded in the next frame image is determined, including: obtaining the current occlusion part of the physical object by the hand in the current frame image; taking the product of the motion speed and the interval between two frames as the occlusion change distance; determining the occlusion change part along the motion direction based on the shape of the speaker's hand and the occlusion change distance; and determining the part of the physical object that is occluded in the next frame image based on the occlusion change part and the current occlusion part.
[0015] To improve the solution, during the display of physical objects, the conference camera is controlled to adjust the image to keep the physical object centered in the frame. This includes: obtaining the proportion of the physical object in the frame and a third correspondence; the third correspondence includes a one-to-one correspondence between multiple proportion ranges and multiple focal length adjustment coefficients; using the focal length adjustment coefficient corresponding to the proportion range of the physical object in the frame in the third correspondence as the target focal length adjustment coefficient; multiplying the current focal length of the conference camera and the target focal length adjustment coefficient as the adjusted focal length of the conference camera; identifying marker points on the physical object; and controlling the conference camera to capture the physical object at the adjusted focal length and tracking the marker points on the physical object to keep the physical object centered in the frame.
[0016] Secondly, this application also discloses a target detection system for a conference camera based on image recognition, comprising: an acquisition device and a processing device; the acquisition device is used to acquire a behavioral image of a speaker; the behavioral image includes a gesture direction and a gaze direction; the gesture direction is the direction pointed by the speaker's index finger; the processing device is used to determine whether the speaker intends to demonstrate a physical object based on the behavioral image, and when the speaker intends to demonstrate a physical object, control the conference camera to point at the area where the physical object is located; the processing device is used to determine the object attribute information of the physical object based on the original image information of the physical object, and adjust the original image attribute information of the conference camera to obtain adjusted image attribute information; the processing device is used to display the physical object in the display screen based on the adjusted image attribute information according to the relative motion relationship between the speaker's hand and the physical object during the speaker's operation of the physical object; the processing device is used to control the conference camera to adjust the image during the display of the physical object so that the physical object remains in the center of the screen.
[0017] Beneficial effects
[0018] This application provides a target detection method for a conference camera based on image recognition. It acquires images of the speaker's behavior, including gesture direction and gaze direction, and intelligently determines whether the speaker intends to demonstrate a physical object based on these images. When a demonstration is detected, the system can precisely control the conference camera to focus on the area where the physical object is located. Furthermore, this application determines the object's attribute information based on the original image information of the physical object and adjusts the original image attribute information of the conference camera to obtain a better image display effect. During the speaker's manipulation of the physical object, this application can clearly display the physical object on the display screen based on the relative movement relationship between the speaker's hand and the physical object, using the adjusted image attribute information. In addition, during the display of the physical object, this application can also control the conference camera to adjust the image, ensuring that the physical object remains centered in the frame.
[0019] Through the above technical solutions, this application effectively solves the technical problems of insufficient recognition and tracking capabilities, unstable image display, and difficulty for remote participants to see the details of physical objects when processing physical object demonstrations in existing intelligent conference camera systems. Specifically, this application, through comprehensive analysis of gesture direction and gaze direction, can more accurately determine the speaker's presentation diagram, avoiding the limitations of traditional systems that rely solely on face or body recognition. Simultaneously, dynamic adjustment of camera image attributes (such as exposure value and frame rate) ensures that physical objects present optimal visual effects under different lighting and motion conditions. Especially when a physical object is obscured by a hand, this application can predict the obscured area based on relative motion and perform image filling, greatly improving the continuity and integrity of the image. Finally, through focus adjustment and marker tracking, the physical object is always kept in the center of the image, providing remote participants with a stable, clear, and uninterrupted viewing experience, significantly improving the effectiveness and participation of physical object demonstrations in meetings, overcoming the shortcomings of existing technologies such as image jitter, target loss, and blurred details, achieving significant technical progress and unexpected technical effects. Attached Figure Description
[0020] Figure 1 A flowchart illustrating a target detection method for a conference camera based on image recognition provided in this application;
[0021] Figure 2 A flowchart illustrating another image recognition-based target detection method for conference cameras provided in this application;
[0022] Figure 3 This application provides a schematic diagram of the architecture of a conference camera target detection system based on image recognition. Detailed Implementation
[0023] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0024] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0025] In modern meeting environments, intelligent meeting camera systems are widely used to automatically detect and accurately track key targets in meetings, such as speakers or presentation devices, through image recognition technology, thereby enhancing the experience for remote participants. However, existing intelligent meeting camera systems face significant technical challenges when handling the presentation of physical objects in meetings. Specifically, in some professional meetings, the focus of discussion may not be limited to human-screen interaction but may also involve the display and discussion of physical objects. When faced with a relatively small, irregularly shaped, and frequently manipulated physical object, the system's ability to identify and continuously track it becomes insufficient. It may fail to accurately identify the physical object as a key target, or even if it does, it may struggle to track it as steadily as it tracks a face. The system may mistakenly focus on the presenter's hand or simply an area of the table, failing to clearly capture the key details of the physical object being discussed. This prevents remote participants from clearly seeing the object's specific features, severely impacting their engagement in the product discussion.
[0026] In this regard, such as Figure 1 As shown, this application proposes a target detection method for conference cameras based on image recognition, including:
[0027] S101. Obtain a behavioral image of the speaker; the behavioral image includes the direction of the gesture and the direction of the gaze; the direction of the gesture is the direction the speaker's index finger is pointing.
[0028] S102. Determine whether the speaker intends to demonstrate a physical object based on the behavioral image, and when the speaker intends to demonstrate a physical object, control the conference camera to point at the area where the physical object is located.
[0029] S103. Determine the object attribute information of the physical object based on the original image information of the physical object, adjust the original image attribute information of the conference camera, and obtain the adjusted image attribute information.
[0030] S104. During the process of the speaker operating the physical object, based on the relative motion relationship between the speaker's hand and the physical object, the physical object is displayed on the display screen based on the adjusted screen attribute information.
[0031] S105. During the display of the physical object, control the conference camera to adjust the image so that the physical object remains in the center of the image.
[0032] This application acquires images of the speaker's behavior and determines whether the speaker intends to demonstrate a physical object based on these images, thereby controlling the conference camera to focus on the area where the physical object is located. Furthermore, by adjusting the conference camera's image attributes and, during the speaker's manipulation of the physical object, based on the relative movement between the speaker's hand and the physical object, the physical object is clearly displayed on the screen, while the conference camera is controlled to adjust the image to keep the physical object centered in the frame. Therefore, this application effectively solves the problems of inaccurate recognition and tracking and unclear image display in existing technologies when conference cameras are used to demonstrate physical objects, significantly improving the meeting experience for remote participants.
[0033] To better understand the technical solution proposed in this application, some key terms and implementation environments involved are first explained. In this application, "behavioral image" refers to the visual information of a speaker captured through image recognition technology, its core being the inclusion of gesture direction and gaze direction. The gesture direction specifically refers to the direction the speaker's index finger points, a clear and easily identifiable indicative action used to express the speaker's attention to or instruction of a specific physical object. The gaze direction reflects the area the speaker's gaze is focused on; combined with the gesture direction, it allows for a more accurate judgment of the speaker's intention. The implementation environment of this application is typically a conference room system equipped with a smart conference camera, display device, and image processing unit. The conference camera is responsible for capturing the video stream of the conference scene, while the image processing unit analyzes the video stream in real time, identifies the speaker's behavioral image and physical objects, and controls the actions of the conference camera and the presentation of the displayed image based on the recognition results.
[0034] The image recognition-based target detection method for conference cameras proposed in this application is based on a series of collaborative technical features that enable accurate identification, tracking, and optimized display of physical objects in a conference.
[0035] First, regarding the feature of acquiring a speaker's behavioral image, the purpose is to capture the speaker's intention signals in interacting with physical objects. One implementation involves a conference camera continuously capturing a video stream from the conference room and using a built-in image recognition module to analyze the people in the video stream in real time. This module can recognize the speaker's body posture, hand gestures, and head orientation. For example, when a speaker raises their arm and points their index finger in a certain direction, the system recognizes the gesture and calculates the vector direction in which the index finger is pointing, i.e., the gesture direction. Simultaneously, through facial recognition and eye-tracking technology, the system can determine the speaker's gaze direction, i.e., the direction of their gaze. This directional information is extracted in real time and serves as a key component of the behavioral image.
[0036] Secondly, regarding the feature of determining whether a speaker intends to demonstrate a physical object based on behavioral images, and controlling the conference camera to focus on the area where the physical object is located when the speaker intends to demonstrate a physical object, the purpose is to quickly shift the camera focus to the target physical object when the speaker is identified as having a demonstration intention. One implementation method is that the system continuously analyzes the acquired gesture direction and gaze direction. When the gesture direction and gaze direction intersect in space, and this intersection is within the current field of view of the conference camera, the system will initially determine that the speaker may be demonstrating a physical object. At this time, the system will further perform image recognition in the intersection area to determine whether there is a non-human entity that can be identified as a physical object. If a physical object is identified, it is determined that the speaker intends to demonstrate a physical object, and the area where the physical object is located is designated as the target area. The conference camera is then controlled to pan, tilt, or zoom, so that its center of view is focused on the target area.
[0037] Secondly, regarding the feature of determining the object attribute information of a physical object based on its original image information, and adjusting the original image attribute information of the conference camera to obtain adjusted image attribute information, the purpose is to optimize the display effect of the physical object. One implementation method is that once a physical object is identified and aligned, the system immediately acquires the original image information of that physical object, such as its color, brightness, and texture. Based on this information, the system can determine the object attribute information of the physical object, such as its dominant color and surface reflectivity. Simultaneously, the system acquires the current original image attribute information of the conference camera, such as the original exposure value and original frame rate. Then, the system makes targeted adjustments to the original image attribute information of the conference camera based on the object attribute information of the physical object. For example, if the physical object is darker, the system may increase the exposure value; if the physical object is moving faster, the system may increase the frame rate to reduce motion blur, thereby obtaining adjusted image attribute information.
[0038] Next, regarding the feature of displaying the physical object on the screen based on the relative movement between the speaker's hand and the physical object, and using adjusted image attribute information, the aim is to ensure the complete display of the physical object even if the speaker's hand obscures it. One implementation method is for the system to continuously monitor the relative position and movement between the speaker's hand and the physical object. When the system detects that the speaker's hand begins to obscure the physical object, it predicts the specific part of the physical object that will be obscured in the next frame based on the movement trajectory, speed, and direction of the hand and the physical object. Then, the system can use the pre-collected complete image data of the physical object to intelligently fill or repair the obscured part in the next frame, combining this with previously adjusted image attribute information (such as exposure, frame rate, etc.) to ensure that the physical object is always presented completely and clearly on the screen.
[0039] Finally, regarding the feature of controlling the conference camera to adjust the image during the display of a physical object, ensuring the object remains centered in the frame, the aim is to guarantee that the object is always in focus. One implementation method is for the system to analyze the position and size of the physical object in the current display frame in real time. If the physical object deviates from the center of the frame or its proportion in the frame is unsatisfactory, the system calculates the necessary image adjustment. For example, the system can dynamically adjust the focal length of the conference camera based on the proportion of the physical object in the frame, making it appear at an appropriate size. Simultaneously, the system identifies one or more landmarks on the physical object (e.g., the geometric center of the object or a specific feature point) and continuously tracks these landmarks. By controlling the pan, tilt, and zoom of the conference camera, these landmarks are kept in the central area of the display frame, thus ensuring the physical object remains centered in the image.
[0040] The image recognition-based target detection method for conference cameras proposed in this application, through the close cooperation of the above-mentioned technical features, can effectively solve many problems existing in the prior art of conference cameras in physical object demonstration scenarios.
[0041] Specifically, traditional smart conference camera systems often face challenges in accurately recognizing and tracking physical objects during presentations, resulting in unclear image display. For example, when a speaker demonstrates a small, irregularly shaped object that may be frequently manipulated, existing systems may fail to accurately identify it as a key target, or even if they do, they may struggle to track it as steadily as they would a face. The system might mistakenly focus on the presenter's hand or simply a certain area of the table, failing to clearly capture the crucial details of the object being discussed. This prevents remote participants from clearly seeing the object's specific features, severely impacting their engagement in the product discussion.
[0042] This application proactively captures the speaker's presentation by introducing a step of obtaining behavioral images of the speaker, rather than passively waiting for the physical object to appear. By combining gesture direction and gaze direction, the system can more accurately determine the area of focus of the speaker's attention, thus controlling the conference camera to point at the area where the physical object is located when the speaker intends to demonstrate it. This proactive recognition and alignment mechanism significantly improves the accuracy and response speed of target detection, avoiding the image clutter caused by unclear targets in existing systems.
[0043] Furthermore, this application determines the object attribute information of a physical object based on its original image information, adjusts the original image attribute information of the conference camera to obtain adjusted image attribute information, and solves the problem of poor display effect of physical objects. For example, for physical objects of different colors, materials, or moving speeds, the system can intelligently adjust parameters such as camera exposure value and frame rate to ensure that the physical object presents the best visual effect in the display screen, avoiding the situation in existing systems where the image parameters are too simple, resulting in some physical objects being displayed too dark, too bright, or blurry.
[0044] Furthermore, this application enhances the user experience by optimizing the display of physical objects based on the relative motion between the speaker's hand and the object, adjusting the display based on the adjusted image attributes, and by controlling the conference camera to center the object during display. When the speaker's hand obscures the object, the system predicts the obstruction based on the relative motion and intelligently fills in the obstruction, ensuring the object's integrity. Simultaneously, by tracking the object in real-time and adjusting the camera focus, the system ensures the object remains centered, allowing remote participants to maintain a clear and continuous focus on the presentation's key points.
[0045] In summary, this application overcomes the limitations of existing technologies in physical object demonstration scenarios by introducing a series of innovative techniques, including speaker behavior image analysis, intelligent image attribute adjustment, hand occlusion prediction and filling, and centralized display of physical objects. Compared to existing technologies, this application achieves accurate, stable, and clear display of physical objects, greatly improving the effectiveness of physical object demonstrations in remote meetings and enhancing the participant experience, demonstrating significant technological advancement and practical value.
[0046] This application further proposes the steps of determining whether a speaker intends to demonstrate a physical object based on behavioral images, and controlling the conference camera to point at the area where the physical object is located when the speaker intends to demonstrate a physical object, including:
[0047] The system determines whether the gesture direction and the line of sight intersect, and whether the intersection point is within the field of view of the conference camera. If the gesture direction and the line of sight intersect, and the intersection point is within the field of view of the conference camera, an image recognition algorithm is used to determine whether a physical object is detected at the intersection point. If a physical object is detected at the intersection point, it is determined that the speaker intends to demonstrate the physical object, and the area where the intersection point exists is taken as the area where the physical object is located, and the conference camera is controlled to point at the area where the physical object is located. If the gesture direction and the line of sight do not intersect, or if the gesture direction and the line of sight intersect, and the intersection point is not within the field of view of the conference camera, or if the gesture direction and the line of sight intersect, and the intersection point is within the field of view of the conference camera, and no physical object is detected at the intersection point, it is determined that the speaker does not intend to demonstrate the physical object.
[0048] Specifically, the direction of a gesture refers to the direction the speaker's index finger points, and the direction of their gaze refers to the direction their eyes are looking. By analyzing the speaker's behavioral images, the direction of the gesture and the direction of the gaze can be extracted. Determining whether the gesture and gaze directions intersect aims to initially judge whether the speaker intends to point to a specific area based on their body language and eye focus. Simultaneously, it is necessary to determine whether this intersection point is within the current field of view of the conference camera to ensure that subsequent object detection and camera alignment operations are feasible.
[0049] Specifically, when a gesture intersects with the direction of the gaze, and this intersection is within the field of view of the conference camera, the system further analyzes the area containing the intersection using an image recognition algorithm. This image recognition algorithm can be trained to recognize various common physical objects, such as documents, models, and electronic devices. In this way, it can be confirmed that a physical object that can be demonstrated does indeed exist in the area the speaker is pointing to.
[0050] In practical applications, if the image recognition algorithm successfully identifies a physical object at the intersection, it can be definitively determined that the speaker intends to demonstrate that physical object. At this point, the area containing the intersection is identified as the location of the physical object. Subsequently, the conference camera will be controlled to point at the area containing the physical object, thus bringing the object into the camera's field of view.
[0051] As a specific implementation, it can be determined that the speaker does not intend to demonstrate a physical object when: the gesture direction and the line of sight do not intersect; or, the gesture direction and the line of sight intersect but the intersection point is not within the field of view of the conference camera; or, even if the intersection point is within the field of view but the image recognition algorithm fails to identify the physical object. These conditions aim to eliminate false judgments and ensure that the camera alignment operation is triggered only when the speaker clearly points and there is an identifiable physical object.
[0052] This application's solution, by combining the speaker's gesture direction and gaze direction, and further utilizing image recognition technology, can more accurately determine whether the speaker intends to demonstrate a physical object. First, the cross-verification of gesture and gaze directions simulates the natural human behavior of pointing at an object—pointing at the target while looking at it—providing a basis for initially locating the target area. Second, limiting the intersection point to the conference camera's field of view ensures the feasibility of subsequent operations. Building on this, an image recognition algorithm is introduced to identify the physical object in the intersection area, further verifying the authenticity of the speaker's intention and avoiding misjudgments caused by accidental intersections of gestures or gazes. Through this multi-verification mechanism, the system can reliably identify the physical object the speaker truly intends to demonstrate and promptly adjust the camera to align it with the target area.
[0053] like Figure 2 As shown, this application further proposes the following steps when, upon identifying the presence of a physical object at an intersection, determining that the speaker intends to demonstrate the physical object, designating the area where the intersection exists as the area where the physical object is located, and controlling the conference camera to point at the area where the physical object is located:
[0054] S201. Obtain the speaker's voice information.
[0055] S202. Convert the speech information into text to obtain text information, and determine whether there are preset words in the text information.
[0056] S203. When there are preset words in the text information, determine that the speaker intends to demonstrate a physical object, take the area where the intersection exists as the area where the physical object is located, and control the conference camera to point at the area where the physical object is located.
[0057] Specifically, acquiring a speaker's voice information refers to collecting the speaker's audio data in real time during the presentation using a microphone array deployed in the conference room or a microphone worn by the speaker. This audio data can include information such as the speaker's speech content, tone, and speaking speed. Voice information can be understood as the sound signal containing the speaker's verbal expression, and its purpose is to provide raw data for subsequent text conversion.
[0058] Furthermore, the speech information is converted into text, and the system determines whether preset words are present in the text. Specifically, the speech information is input into the speech recognition module, which uses acoustic and language models to convert continuous speech signals into discrete text sequences, i.e., text information. Preset words refer to a set of keywords or phrases highly relevant to the demonstration or introduction of a physical object, such as "please look," "this," "demonstration," "model," "chart," "everyone pay attention," etc. The system uses natural language processing technology to perform semantic analysis and keyword matching on the converted text information to determine whether it contains these preset words. The purpose is to help determine the speaker's true intention through language content.
[0059] In practical applications, when preset words are present in the text information, the system will further confirm the speaker's intention to demonstrate a physical object. This means that even if there is some uncertainty in the image recognition results, combining the speaker's clear verbal expression can more reliably determine the demonstration. At this point, the system will identify the area where the gesture direction and the line of sight intersect as the area where the physical object is located, and immediately control the conference camera to point at that area in order to clearly capture and display the physical object.
[0060] This application's solution effectively overcomes the limitations of relying solely on image recognition when determining a speaker's intentions by introducing analysis of the speaker's voice information. When the image recognition algorithm identifies a physical object at an intersection, the system does not immediately determine the presence of a speaker's intention; instead, it further acquires and analyzes the speaker's voice information. Because a speaker's language often reflects their intentions more directly and clearly, by converting the voice information into text and determining whether pre-defined words are present, the system can perform secondary verification or supplementary confirmation of the image recognition results. For example, when a speaker points to an object and simultaneously says "Please look at this chart," the pre-defined words "please look" and "chart" in the voice information strongly corroborate the intention, thus avoiding misjudgments that might arise from purely visual judgment. This multimodal information fusion judgment mechanism enables the system to understand the speaker's true intentions more accurately and intelligently in complex and ever-changing meeting environments.
[0061] This application further proposes a method for fine-tuning the attribute information of conference camera images based on the characteristics of physical objects.
[0062] Object attribute information includes the physical object's color and movement speed; original image attribute information includes the original exposure value and original frame rate. Based on the original image information of the physical object, the object attribute information is determined. The original image attribute information of the conference camera is adjusted to obtain the adjusted image attribute information, including:
[0063] The color of the physical object is determined based on its original image information, and its hue value is determined based on the color of the physical object. The product of the hue value and the preset color adjustment coefficient is used as the exposure compensation gain. The original exposure value of the conference camera is adjusted based on the exposure compensation gain to obtain the adjusted exposure value. The moving speed of the physical object is determined based on its original image information, and the original frame rate of the conference camera is adjusted based on the moving speed of the physical object to obtain the adjusted frame rate.
[0064] Specifically, object attribute information refers to the inherent visual or motion characteristics of a physical object, such as its color and speed of movement. The color of a physical object can be understood as the characteristic of light reflected from its surface, which can be extracted from the original image information of the physical object using image recognition technology. Hue value is a quantitative representation of color, used to describe the type of color, such as red, blue, etc. The preset color adjustment coefficient is a value pre-set based on experience or experimentation, used to adjust the degree of influence of the hue value on the exposure compensation gain, with the aim of ensuring that physical objects of different colors achieve appropriate brightness in the image. Exposure compensation gain is a correction amount used to adjust the camera's exposure value, its purpose being to compensate for brightness deviations caused by the color characteristics of the physical object.
[0065] Raw image attribute information refers to the default or current image parameters of the conference camera when no specific adjustments have been made, such as the raw exposure value and the raw frame rate. The raw exposure value determines the overall brightness of the image, while the raw frame rate affects the smoothness of the image and the ability to capture moving objects.
[0066] In practical applications, the movement speed of a physical object can be determined based on its original image information by calculating the changes in its position across consecutive frames. For example, image processing techniques such as optical flow and feature point tracking can be used to estimate the object's trajectory and speed within the frame. Adjusting the frame rate of the conference camera's original footage based on the object's movement speed aims to ensure the image remains clear and smooth even when the object is moving rapidly, reducing motion blur. For instance, when the object moves quickly, the frame rate can be increased to capture more details; conversely, when the object moves slowly, the frame rate can be decreased to conserve processing resources.
[0067] This application's solution achieves fine-tuning of conference camera image attributes by identifying and analyzing the color and movement speed of physical objects. Specifically, by determining the color of the physical object and calculating its hue value, and combining this with a preset color adjustment coefficient to generate exposure compensation gain, the original exposure value is adjusted. This effectively solves the problem of uneven image brightness caused by differences in the color of physical objects, ensuring that physical objects present accurate color and brightness in the image. Simultaneously, by monitoring the movement speed of physical objects in real time and adjusting the original frame rate of the conference camera accordingly, motion blur that may occur when physical objects move rapidly can be effectively addressed, ensuring image clarity and smoothness. It is precisely this targeted adjustment of image attributes based on the characteristics of physical objects that allows the conference camera to better adapt to different presentation scenarios and physical objects, providing high-quality visual presentations.
[0068] This application further proposes the following steps for adjusting the original exposure value of the conference camera based on the exposure compensation gain to obtain the adjusted exposure value:
[0069] The light intensity of the conference room detected by the light sensor is obtained and a first correspondence is established. The first correspondence includes a one-to-one correspondence between multiple light intensity ranges and multiple light adjustment coefficients. The light adjustment coefficient corresponding to the light intensity range in which the light intensity of the conference room is located is taken as the target light adjustment coefficient. The product of the exposure compensation gain and the target light adjustment coefficient is taken as the target exposure compensation gain. The sum of the original exposure value of the conference camera and the target exposure compensation gain is taken as the adjusted exposure value of the conference camera.
[0070] Specifically, when adjusting the raw exposure value of the conference camera, the current light intensity of the conference room needs to be obtained first. This light intensity can be detected in real time by light sensors deployed in the conference room. Simultaneously, the system presets and stores a first correspondence, which defines in detail the mapping relationship between different light intensity ranges and their corresponding light adjustment coefficients. For example, when the light intensity is in the range of "0-100 lux", the corresponding light adjustment coefficient may be 1.2; when the light intensity is in the range of "101-500 lux", the corresponding light adjustment coefficient may be 1.0; when the light intensity is in the range of "501-1000 lux", the corresponding light adjustment coefficient may be 0.8, and so on.
[0071] After obtaining the current light intensity in the conference room, the system will use a first correspondence to find the specific range of that light intensity and extract the corresponding light adjustment coefficient, which will be used as the target light adjustment coefficient. For example, if the current light intensity in the conference room is 300 lux, it will match the range of "101-500 lux" and obtain the corresponding light adjustment coefficient of 1.0 as the target light adjustment coefficient.
[0072] In practical applications, to comprehensively consider the color characteristics of the physical object and the overall lighting environment of the conference room, the exposure compensation gain previously determined based on the color of the physical object is multiplied by the target light adjustment coefficient obtained above, resulting in a more refined target exposure compensation gain. This target exposure compensation gain not only reflects the exposure requirements of the physical object itself but also incorporates the influence of ambient light on exposure.
[0073] Finally, the original exposure value of the conference camera is summed with the target exposure compensation gain, and the result is the final adjusted exposure value of the conference camera. This adjusted exposure value can more accurately adapt to the current lighting environment of the conference room, ensuring that physical objects receive the best exposure effect on the display screen.
[0074] This application's solution incorporates the ambient light intensity detected by a light sensor in the conference room and combines it with a preset first correspondence to obtain a target light adjustment coefficient, thereby correcting the exposure compensation gain determined based on the physical object's color. Specifically, when the conference room is dimly lit, the target light adjustment coefficient may be set to a value greater than 1, increasing the final target exposure compensation gain and thus increasing the exposure, preventing the physical object from being underexposed in low-light environments. Conversely, when the conference room is brightly lit, the target light adjustment coefficient may be set to a value less than 1, decreasing the final target exposure compensation gain and thus reducing the exposure, preventing the physical object from being overexposed in bright light environments. Therefore, by incorporating ambient light factors into the exposure adjustment, the final exposure value better adapts to the actual environment, solving the problem of inaccurate exposure that may occur when adjusting exposure solely based on the physical object's color.
[0075] This application further proposes a method for adjusting the frame rate of the raw footage from a conference camera based on the movement speed of physical objects, including:
[0076] Obtain the second correspondence and the importance index of the current meeting; the second correspondence includes a one-to-one correspondence between multiple movement speed ranges and multiple frame rate adjustment coefficients; use the multiple frame rate adjustment coefficients corresponding to the movement speed range of the physical object in the second correspondence as the initial frame rate adjustment coefficients; use the weighted sum of the initial frame rate adjustment coefficients and the importance index of the current meeting as the target frame rate adjustment coefficient; use the product of the original frame rate of the conference camera and the target frame rate adjustment coefficient as the adjusted frame rate of the conference camera.
[0077] Specifically, the "second correspondence" refers to a pre-established mapping rule used to guide frame rate adjustments. This correspondence associates different "movement speed ranges" of physical objects with corresponding "frame rate adjustment coefficients." For example, when a physical object moves slowly, the corresponding frame rate adjustment coefficient may be smaller to save resources; when it moves quickly, the corresponding frame rate adjustment coefficient may be larger to ensure smooth gameplay.
[0078] The "Importance Index of the Current Meeting" can be understood as a quantitative indicator measuring the importance of the current meeting. This index can be determined based on various factors, such as the number of participants, the topic of the meeting, the level of the meeting, and the duration of the meeting. For example, a meeting involving major company decisions may have a higher importance index, while a routine departmental meeting may have a lower importance index. The introduction of this index aims to allow frame rate adjustments to better adapt to the actual needs of the meeting.
[0079] The "initial frame rate adjustment coefficient" is a preliminary adjustment coefficient found from the second correspondence based on the movement speed of physical objects.
[0080] The "target frame rate adjustment coefficient" is the final adjustment coefficient obtained by weighting the initial frame rate adjustment coefficient with the current important index. The weighted sum calculation method can be set according to actual needs. For example, different weights can be assigned to the initial frame rate adjustment coefficient and the important index to reflect their relative importance in frame rate adjustment.
[0081] Ultimately, the "adjusted frame rate of the conference camera" is obtained by multiplying the original frame rate of the conference camera by the target frame rate adjustment factor.
[0082] This application's solution allows for more refined adjustment of the conference camera's frame rate by introducing a second correspondence and an importance index for the current meeting. Specifically, the second correspondence provides a basic frame rate adjustment suggestion based on the movement speed of physical objects, ensuring basic smoothness of the footage at different movement speeds. Furthermore, the introduction of the importance index for the current meeting means that frame rate adjustment no longer relies solely on the motion state of physical objects, but also takes into account the contextual information of the meeting itself.
[0083] By weighting the initial frame rate adjustment coefficient with an importance index, a target frame rate adjustment coefficient that comprehensively considers both motion speed and meeting importance can be generated. For example, in important meetings, even if the physical objects are moving at a moderate speed, the frame rate may be appropriately increased by increasing the weight of the importance index to ensure clear presentation of key information; while in less important meetings, the frame rate may be appropriately reduced to conserve system resources while maintaining basic smoothness. This comprehensive adjustment mechanism allows the meeting camera's frame rate to more intelligently adapt to different application scenarios, thereby improving the user experience.
[0084] Through the above technical solution, this application enables adaptive adjustment of the frame rate of the conference camera, taking into account not only the movement speed of physical objects but also the crucial factor of the importance of the meeting. This avoids image stuttering or blurring due to insufficient frame rate in important meetings, thus preventing disruption to information transmission efficiency, while also preventing excessive consumption of system resources in less important meetings. This refined frame rate adjustment mechanism allows the conference camera to provide smoother and clearer images according to actual needs, significantly improving the viewing experience and information delivery effectiveness of meetings.
[0085] This application further proposes a step for displaying a physical object on a screen based on the relative motion between the speaker's hand and the physical object, and according to adjusted screen attribute information, during the speaker's manipulation of the physical object:
[0086] Determine whether the speaker's hand is obscuring the physical object; when the speaker's hand obscures the physical object, determine the obscured part of the physical object in the next frame image based on the relative motion relationship between the speaker's hand and the physical object; fill in the obscured part of the physical object in the next frame image based on the pre-captured image of the physical object, so as to display the physical object in the display screen based on the adjusted screen attribute information.
[0087] Specifically, determining whether a speaker's hand is obscuring a physical object can be achieved using image recognition technology. For example, a deep learning model can be used to analyze real-time footage captured by a conference camera to identify the speaker's hand area and the physical object area. When the hand area and the physical object area overlap in the image, and the overlap reaches a preset threshold, it can be determined that the hand is obscuring the physical object. The preset threshold can be set according to the actual application scenario and the tolerance for the degree of obstruction; for example, when the hand obscures more than 10% of the physical object's area, obstruction is considered to have occurred.
[0088] Furthermore, when the speaker's hand occludes a physical object, the occluded portion of the object in the next frame is determined based on the relative motion between the speaker's hand and the physical object. This aims to predict the occlusion of the physical object by the hand in the next moment. The relative motion relationship can include the speed and direction of the hand's movement relative to the physical object. By analyzing the position, shape, and trajectory of the hand and the physical object in the current frame, the specific area of the physical object that the hand may occlude in the next frame can be inferred.
[0089] Specifically, based on pre-captured images of the physical object, the occluded portions of the physical object in the next frame are filled in. The pre-captured images of the physical object refer to the complete image data of the physical object that the system has acquired and stored before the demonstration begins or when the physical object is not occluded. When the system predicts that the physical object will be occluded by a hand in the next frame, it can extract the image information of the corresponding occluded portion from the pre-captured complete image and seamlessly fill it into the live screen, thus presenting the complete physical object on the display.
[0090] This application's solution effectively addresses the issue of incomplete display caused by a speaker's hand obscuring a physical object. Specifically, by introducing a mechanism for detecting, predicting, and filling in the obscuring of the physical object by the speaker's hand, the system can promptly identify potential display defects. Secondly, upon detecting obscuring, based on the relative motion between the hand and the physical object, the system can accurately predict the specific part of the physical object obscured in the next frame, providing accurate positioning information for subsequent image restoration. Finally, by filling in the obscured area using a pre-captured complete image of the physical object, the system ensures that the physical object in the display maintains its integrity and continuity even with hand obscuring, significantly improving the viewing experience.
[0091] Through the aforementioned technical solution, this application can intelligently identify and handle situations where the speaker's hands obstruct the physical objects being presented, avoiding incomplete images or missing information caused by hand occlusion. This ensures that the physical objects are always presented in a complete and clear form on the display screen throughout the entire presentation, significantly enhancing the professionalism of the meeting and the audience's immersion. Furthermore, this solution achieves seamless real-time image restoration through prediction and infill technology, enabling the audience to continuously receive high-quality visual information and effectively guaranteeing the complete delivery of the presentation content.
[0092] This application further proposes the following steps for determining the occluded portion of a physical object in the next frame image:
[0093] Obtain the current occlusion part of the physical object by the hand in the current frame image; take the product of the movement speed and the time interval between two frames as the occlusion change distance; determine the occlusion change part along the movement direction based on the shape of the speaker's hand and the occlusion change distance; determine the occlusion part of the physical object in the next frame image based on the occlusion change part and the current occlusion part.
[0094] Specifically, the aforementioned relative motion relationship can be understood as the motion state of the speaker's hand relative to a physical object, including motion speed and direction. Motion speed refers to the distance the hand moves per unit time, while motion direction refers to the trajectory direction of the hand's movement. These motion parameters can be obtained by analyzing and calculating the positional changes of the hand and the physical object in consecutive frames using image processing techniques such as optical flow, feature point tracking, or deep learning models.
[0095] Specifically, obtaining the current occlusion area of the hand on the physical object in the current frame image refers to identifying the area where the hand overlaps with the physical object in the current frame using techniques such as image segmentation and object detection. This area is the part of the physical object that is occluded by the hand in the current frame.
[0096] Furthermore, the product of the movement speed and the interval between two frames is taken as the occlusion change distance. The interval between two frames refers to the time interval between the system processing two consecutive frames, which is usually a fixed value. By multiplying the hand's movement speed by this interval, the displacement of the hand relative to its current position in the next frame can be predicted, i.e., the occlusion change distance.
[0097] Based on this, the occlusion change location is determined along the direction of movement, according to the shape of the speaker's hand and the occlusion change distance. The shape of the speaker's hand can be obtained through a pre-trained model or real-time image analysis. Combining the direction of hand movement and the predicted occlusion change distance, the new area that the hand may cover in the next frame image can be inferred, i.e., the occlusion change location. For example, if the hand moves to the right, the occlusion change location will be the newly added occluded area to the right of the current occluded area.
[0098] Finally, based on the occlusion change location and the current occlusion location, the occluded portion of the physical object in the next frame is determined. This typically means logically combining the occluded portion of the current frame with the predicted occlusion change location (e.g., taking the union) to obtain the complete area of the physical object occluded by the hand in the next frame.
[0099] This application's solution refines the relative motion relationship into motion speed and direction, and combines this with the current occlusion location, the predicted occlusion change distance, and the hand shape to more accurately predict the specific area of a physical object occluded by a hand in the next frame. Traditional methods may rely solely on simple positional differences for prediction, which is prone to errors when the hand moves rapidly or has a complex shape. This solution, however, obtains precise occlusion information from the current frame, predicts the hand's displacement using kinematic principles, and then infers the region based on the hand shape, resulting in a more refined and accurate determination of the occlusion location in the next frame. It is precisely this refined prediction mechanism that allows subsequent image filling to be more seamless and natural, effectively avoiding image jitter or discontinuity caused by inaccurate predictions.
[0100] The above technical solution significantly improves the accuracy of predicting the obscured portion of a physical object when the speaker's hand is covering it. This precise prediction capability makes subsequent image filling of the obscured area more natural and smooth, effectively reducing visual artifacts and discontinuities in the image. Compared to methods that rely solely on rough relative motion relationships for prediction, this solution better adapts to the complex and varied movements of the speaker's hand, thereby improving the image quality and user experience of the conference camera during the presentation of physical objects, ensuring the integrity and clarity of the physical object in the displayed image.
[0101] This application further proposes a method for controlling a conference camera to adjust the image during the display of a physical object, so that the physical object remains centered in the image, comprising:
[0102] Obtain the proportion of the physical object in the frame and the third correspondence relationship; the third correspondence relationship includes a one-to-one correspondence between multiple proportion ranges and multiple focal length adjustment coefficients; take the focal length adjustment coefficient corresponding to the proportion range of the physical object in the frame in the third correspondence relationship as the target focal length adjustment coefficient; take the product of the current focal length of the conference camera and the target focal length adjustment coefficient as the adjusted focal length of the conference camera; determine the marker point on the physical object; control the conference camera to shoot the physical object with the adjusted focal length and track the marker point on the physical object to keep the physical object in the center of the frame.
[0103] Specifically, in the process of displaying a physical object, to keep the object centered in the frame, it is first necessary to obtain the object's proportion within the frame. This proportion can be understood as the ratio of the pixel area occupied by the physical object in the current camera frame to the total pixel area of the frame, reflecting the object's size and distance information within the frame. Simultaneously, a third correspondence is obtained, which pre-stores a one-to-one correspondence between multiple proportion ranges and multiple focal length adjustment coefficients. For example, when the physical object's proportion is too small, it may be necessary to increase the focal length to enlarge the object; when the proportion is too large, it may be necessary to decrease the focal length to shrink the object, ensuring its complete display.
[0104] Specifically, the focal length adjustment coefficient corresponding to the proportion range of the physical object in the image according to the third correspondence is used as the target focal length adjustment coefficient. This means that the system will find the most suitable focal length adjustment coefficient from the preset correspondence based on the actual size of the physical object in the image. For example, if the proportion of the physical object is within a certain range, the focal length adjustment coefficient corresponding to that range will be selected.
[0105] In practical applications, the product of the current focal length and the target focal length adjustment factor of the conference camera is used as the adjusted focal length of the conference camera. In this way, the focal length of the camera can be dynamically adjusted to adapt to changes in the proportion of physical objects in the frame, thereby enabling the magnification or reduction of physical objects to present the desired size in the image.
[0106] Furthermore, it is necessary to identify landmarks on the physical object. These landmarks can be easily identifiable feature points on the physical object, such as corner points, texture feature points, or pre-defined identification markers. These landmarks are used for subsequent tracking to ensure the position of the physical object in the image.
[0107] Finally, the conference camera is controlled to capture the physical object at the adjusted focal length and track markers on the object to keep it centered in the frame. This means the camera not only adjusts the focal length based on the size of the object but also fine-tunes its pan and tilt angles by tracking its markers, ensuring the object remains centered in the frame.
[0108] This application's solution effectively solves the problem of unstable and inaccurate centering of physical objects in the basic solution by introducing a correspondence between the proportion of physical objects in the frame and the focal length adjustment coefficient, combined with the tracking of physical object markers. Because the focal length is dynamically adjusted based on the proportion of the physical object in the frame, the size of the physical object in the frame is kept within a suitable range, avoiding poor display due to it being too large or too small. Simultaneously, by identifying and tracking markers on the physical object, the conference camera can perceive changes in the object's position in real time and make precise translation and tilt adjustments, ensuring that the physical object is always locked in the center of the frame. This combination of focal length adjustment and position tracking allows the conference camera to intelligently adapt to changes in the size and position of physical objects during the speaker's presentation, providing a stable and centered display effect.
[0109] Through the aforementioned technical solution, this application enables precise control over the adjustment of the conference camera image, allowing physical objects to remain more stably and accurately centered in the frame during display. Compared to simply making general image adjustments, this solution significantly improves the centering stability and visual effect of physical objects in the conference display by introducing a linkage mechanism between the proportion of physical objects and focal length adjustment, as well as precise tracking of physical object markers. This not only optimizes the speaker's presentation experience, eliminating concerns about physical objects deviating from the center of the frame, but also greatly improves the viewing experience for attendees, ensuring clear and focused presentation of the content.
[0110] This application proposes a target detection system for a conference camera based on image recognition, comprising: an acquisition device and a processing device; the acquisition device is used to acquire a behavioral image of a speaker; the behavioral image includes the direction of a gesture and the direction of a gaze; the direction of the gesture is the direction pointed by the speaker's index finger; the processing device is used to determine whether the speaker intends to demonstrate a physical object based on the behavioral image, and when the speaker intends to demonstrate a physical object, control the conference camera to point at the area where the physical object is located; the processing device is used to determine the object attribute information of the physical object based on the original image information of the physical object, and adjust the original image attribute information of the conference camera to obtain adjusted image attribute information; the processing device is used to display the physical object in the display screen based on the adjusted image attribute information according to the relative motion relationship between the speaker's hand and the physical object during the speaker's operation of the physical object; the processing device is used to control the conference camera to adjust the image during the display of the physical object so that the physical object remains in the center of the screen.
[0111] It should be emphasized that the acquisition device can be one or more image sensors, such as the conference camera itself or a standalone vision sensor, which is equipped with an image acquisition module responsible for capturing the video stream in the conference room in real time and transmitting the raw video data to the processing device.
[0112] The processing device can consist of one or more processors, memory, and corresponding software modules. For example, the processing device can be a high-performance embedded system, a server, or a personal computer, running image recognition algorithms, intent determination logic, image attribute adjustment algorithms, occlusion handling algorithms, and a camera control module. The processing device receives image data transmitted from the acquisition device and performs in-depth analysis on it.
[0113] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A target detection method for a conference camera based on image recognition, characterized in that, include: Obtain images of the speaker's behavior; Behavioral images include gesture direction and gaze direction; The direction of the gesture is the direction the speaker's index finger is pointing; Based on the behavioral images, determine whether the speaker intends to demonstrate a physical object, and when the speaker does intend to demonstrate a physical object, control the conference camera to point at the area where the physical object is located; Based on the original image information of the physical object, the object attribute information of the physical object is determined, and the original image attribute information of the conference camera is adjusted to obtain the adjusted image attribute information. During the process of the speaker manipulating the physical object, the physical object is displayed on the screen based on the relative motion relationship between the speaker's hand and the physical object, and on the adjusted screen attribute information. During the display of physical objects, the conference camera is controlled to adjust the image so that the physical objects remain in the center of the frame.
2. The method for target detection in a conference camera based on image recognition according to claim 1, characterized in that, Based on behavioral images, determine whether the speaker intends to demonstrate a physical object, and when the speaker does intend to demonstrate a physical object, control the conference camera to point at the area where the physical object is located, including: Determine whether the direction of the gesture intersects with the direction of the gaze, and whether the point of intersection is within the field of view of the conference camera; When the direction of the gesture intersects with the direction of the gaze, and the intersection point is within the field of view of the conference camera, the image recognition algorithm determines whether a physical object is detected at the intersection point. When a physical object is detected at the intersection, it is determined that the speaker intends to demonstrate the physical object. The area where the intersection exists is taken as the area where the physical object is located, and the conference camera is controlled to point at the area where the physical object is located. When the direction of the gesture does not intersect with the direction of the gaze, or when the direction of the gesture intersects with the direction of the gaze but the intersection point is not within the field of view of the conference camera, or when the direction of the gesture intersects with the direction of the gaze and the intersection point is within the field of view of the conference camera but no physical object is detected at the intersection point, it is determined that the speaker does not intend to demonstrate a physical object.
3. The method for target detection in a conference camera based on image recognition according to claim 2, characterized in that, Upon detecting a physical object at the intersection, it is determined that the speaker intends to demonstrate the physical object. The area where the intersection exists is designated as the area where the physical object is located, and the conference camera is controlled to point at the area where the physical object is located, including: Obtain the speaker's voice information; The speech information is converted into text, and the presence of preset words in the text information is determined. When preset words are present in the text information, it is determined that the speaker intends to demonstrate a physical object. The area where the intersection exists is taken as the area where the physical object is located, and the conference camera is controlled to point at the area where the physical object is located.
4. The method for target detection in a conference camera based on image recognition according to claim 1, characterized in that, Object attribute information includes the physical object's color and movement speed; original image attribute information includes the original exposure value and original frame rate. Based on the original image information of the physical object, the object attribute information is determined. The original image attribute information of the conference camera is adjusted to obtain the adjusted image attribute information, including: The color of the physical object is determined based on its original image information, and the hue value of the physical object is determined based on its color. The product of the hue value and the preset color adjustment coefficient is used as the exposure compensation gain; The original exposure value of the conference camera is adjusted based on the exposure compensation gain to obtain the adjusted exposure value; The movement speed of the physical object is determined based on its original image information, and the frame rate of the conference camera is adjusted based on the movement speed of the physical object to obtain the adjusted frame rate.
5. The method for target detection in a conference camera based on image recognition according to claim 4, characterized in that, The original exposure value of the conference camera is adjusted based on the exposure compensation gain to obtain the adjusted exposure value, including: The light intensity of the conference room detected by the light sensor is obtained and a first correspondence is established; the first correspondence includes a one-to-one correspondence between multiple light intensity ranges and multiple light adjustment coefficients; The light adjustment coefficient corresponding to the light brightness range of the conference room is used as the target light adjustment coefficient; The product of the exposure compensation gain and the target light adjustment factor is used as the target exposure compensation gain; The sum of the original exposure value of the conference camera and the target exposure compensation gain is used as the adjusted exposure value of the conference camera.
6. The method for target detection in a conference camera based on image recognition according to claim 4, characterized in that, Adjusting the frame rate of the conference camera's raw footage based on the movement speed of physical objects includes: Obtain the second correspondence and the important index of the current meeting; the second correspondence includes a one-to-one correspondence between multiple movement speed ranges and multiple frame rate adjustment coefficients; Use the multiple frame rate adjustment coefficients corresponding to the movement speed range of the physical object in the second correspondence as the initial frame rate adjustment coefficients; The target frame rate adjustment factor is the weighted sum of the initial frame rate adjustment factor and the important indices of the current meeting. The adjusted frame rate of the conference camera is the product of the original frame rate and the target frame rate adjustment factor.
7. The method for target detection in a conference camera based on image recognition according to claim 1, characterized in that, During the speaker's manipulation of the physical object, based on the relative motion between the speaker's hand and the physical object, and using adjusted screen attribute information, the physical object is displayed on the screen, including: Determine whether the speaker's hand is obscuring a physical object; When the speaker's hand occludes a physical object, the part of the physical object that is occluded in the next frame is determined based on the relative motion between the speaker's hand and the physical object. Based on the pre-acquired images of the physical objects, the occluded parts of the physical objects in the next frame image are filled in, so that the physical objects are displayed on the display screen based on the adjusted screen attribute information.
8. The method for target detection in a conference camera based on image recognition according to claim 7, characterized in that, Relative motion relationships include speed and direction of motion. Based on the relative motion relationship between the speaker's hand and the physical object, the occluded portion of the physical object in the next frame is determined, including: Get the current occlusion part of the hand on the physical object in the current frame image; The product of motion speed and the time interval between two frames is used as the occlusion change distance; Along the direction of movement, determine the location of the occlusion change based on the shape of the speaker's hand and the distance of the occlusion change; Based on the changing occlusion location and the current occlusion location, determine the occluded location of the physical object in the next frame image.
9. The method for target detection in a conference camera based on image recognition according to claim 1, characterized in that, During the display of physical objects, the conference camera is controlled to adjust the image to keep the physical objects centered in the frame, including: Obtain the proportion of physical objects in the image and the third correspondence; the third correspondence includes a one-to-one correspondence between multiple proportion ranges and multiple focal length adjustment coefficients; The focal length adjustment coefficient corresponding to the proportion range of the physical object in the image in the third correspondence is used as the target focal length adjustment coefficient. The product of the current focal length and the target focal length adjustment factor of the conference camera is used as the adjusted focal length of the conference camera. Determine the marker points on the physical object; Control the conference camera to capture the physical object at the adjusted focal length and track markers on the physical object to keep the object centered in the frame.
10. A target detection system for a conference camera based on image recognition, characterized in that, include: Acquisition device and processing device; Acquisition device, used to acquire images of the speaker's behavior; Behavioral images include gesture direction and gaze direction; The direction of the gesture is the direction the speaker's index finger is pointing; The processing device is used to determine whether the speaker intends to demonstrate a physical object based on the behavioral image, and when the speaker intends to demonstrate a physical object, to control the conference camera to point at the area where the physical object is located; The processing device is used to determine the object attribute information of the physical object based on the original image information of the physical object, adjust the original image attribute information of the conference camera, and obtain the adjusted image attribute information. The processing device is used to display the physical object on the display screen based on the relative motion relationship between the speaker's hand and the physical object and the adjusted screen attribute information during the speaker's operation of the physical object. The processing device is used to control the conference camera to adjust the image during the display of a physical object, so that the physical object remains in the center of the image.