Intelligent learning machine vision control system, method and related devices
Patent Information
- Application Number
- CN202611004803.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]然而,现有方案仍存在明显缺点
[0048]借由上述技术方案,本申请提供的智能学习机视觉控制系统、方法及相关装置,利用高度集成、非机械或少机械运动的单颗可变焦光学镜头替代了传统学习机的多摄像头或可旋转机械云台单摄像头,通过标定单元保障了单颗可变焦光学镜头在全生命周期内的物理性能;通过处理器与场景识别机制赋予了单颗可变焦光学镜头应对智能学习机多学习场景任务需求的能力,因此,该方案实现了在保证高性能交互体验的前提下,通过硬件的高度集成降低了物料成本,同时凭借软件层面的场景自适应能力,完美匹配了智能学习机复杂多变的使用场景。
Smart Images

Figure CN122802767A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent learning machine technology, and in particular to an intelligent learning machine vision control system, method and related device. Background Technology
[0002] With the deep integration of "AI + Education," smart learning machines have become a core tool to assist K-12 students in their learning. Currently, most mainstream high-end learning machines are equipped with cameras to meet the visual perception needs of different learning scenarios, such as real-time monitoring of students' posture and attention to protect their eyesight and spinal health; and photographing homework and test papers for intelligent diagnosis and grading. Existing technologies mainly include two solutions: one is a multi-camera fixed-function solution, and the other is a single-camera solution with a rotatable mechanical gimbal.
[0003] However, existing solutions still have significant drawbacks. Multi-camera solutions require the procurement, installation, and calibration of multiple camera modules, resulting in high material costs and assembly complexity, significantly increasing BOM costs. Multiple independent modules cannot operate simultaneously, lacking scene fusion capabilities. Multiple camera holes disrupt the device's integrity, affecting aesthetics and encroaching on internal space, hindering slim and lightweight designs. Rotatable mechanical gimbal solutions, on the other hand, suffer from wear and dust ingress risks in moving mechanical parts, are prone to failure with prolonged use, and are more vulnerable to drops. Gimbal rotation is accompanied by noise and delay (typically 1-2 seconds), preventing instantaneous and seamless scene switching, limiting dynamic performance, and causing long startup response times and image blurring during rotation.
[0004] Therefore, how to provide a highly integrated, cost-controllable, and adaptable visual control solution for intelligent learning machines that can adapt to different learning scenarios has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of the above problems, this application provides a visual control system, method, and related apparatus for intelligent learning machines, aiming to achieve high integration, controllable cost, and adaptability to different learning scenarios of intelligent learning machines. The specific solution is as follows:
[0006] The first aspect of this application provides a vision control system for an intelligent learning machine, comprising:
[0007] Monocular adaptive vision acquisition module and processor;
[0008] The monocular adaptive vision acquisition module includes a variable focal length optical lens and a calibration unit;
[0009] The processor is used to calibrate the variable-focus optical lens using the calibration unit to obtain calibration parameters corresponding to different learning scenarios of the intelligent learning machine; and during system operation, it automatically identifies the current learning scenario and at least one observation target contained in the current learning scenario; formulates a control strategy based on the identification result and controls the variable-focus optical lens to observe the at least one observation target in sequence according to the control strategy, and directly applies the calibration parameters corresponding to the current learning scenario to control the variable-focus optical lens during observation.
[0010] In one possible implementation, the variable-focus optical lens is a liquid lens.
[0011] In one possible implementation, the calibration unit has at least two built-in distance calibration benchmarks, with different distance calibration benchmarks corresponding to different learning scenarios and different calibration reference objects.
[0012] In one possible implementation, an environment interaction unit is also included. In this case, the processor is further configured to identify the current learning scene during system operation and, when directly applying the calibration parameters corresponding to the current learning scene to control the variable focal length optical lens, coordinate with the environment interaction unit to perform environmental adjustments.
[0013] A second aspect of this application provides a visual control method for an intelligent learning machine, applied to a processor in a visual control system for an intelligent learning machine. The intelligent learning machine visual control system is an intelligent learning machine visual control system according to the first aspect or any implementation thereof, comprising:
[0014] The calibration unit is used to calibrate the variable-focus optical lens to obtain calibration parameters corresponding to different learning scenarios of the intelligent learning machine.
[0015] During system operation, the system automatically identifies the current learning scene and at least one observation target contained in the current learning scene; based on the identification results, a control strategy is formulated and the variable focal length optical lens is controlled to observe the at least one observation target in sequence according to the control strategy. During observation, the calibration parameters corresponding to the current learning scene are directly applied to control the variable focal length optical lens.
[0016] In one possible implementation, calibrating the variable-focus optical lens using the calibration unit to obtain calibration parameters of the variable-focus optical lens corresponding to different learning scenarios includes:
[0017] Determine whether the preset calibration trigger conditions are met;
[0018] The variable-focus optical lens is controlled to optimize optical parameters for different distance calibration benchmarks, thereby determining the calibration parameters of the variable-focus optical lens for different learning scenarios of the intelligent learning machine.
[0019] In one possible implementation, during system operation, the system automatically identifies the current learning scene and at least one observation target contained within the current learning scene; based on the identification result, a control strategy is formulated, and the variable-focus optical lens is controlled according to the control strategy to sequentially observe the at least one observation target. During observation, calibration parameters corresponding to the current learning scene are directly applied to control the variable-focus optical lens, including:
[0020] The variable-focus optical lens is controlled to perform active scanning to acquire multiple frames of raw environmental images;
[0021] The multi-frame original environmental images are analyzed to obtain the learning scene recognition result; the learning scene recognition result is used to indicate the current learning scene and at least one observation target contained in the current learning scene and the type corresponding to each observation target;
[0022] A control strategy is formulated based on the learning scene recognition results. The control strategy is used to indicate the observation plan and corresponding lens parameters for each target in the current environment.
[0023] The variable focal length optical lens is controlled according to the control strategy.
[0024] In one possible implementation, the step of formulating a control strategy based on the learned scene recognition result includes:
[0025] Based on the current learning scenario and the type corresponding to each observation target, evaluate the observation priority of each observation target;
[0026] Based on the observation priority, the time sequence for observing each of the observation targets and the corresponding lens parameters are generated.
[0027] When the observation time windows of at least two of the observation targets overlap, conflict resolution is performed according to the observation priority to determine the actual observation time of each of the observation targets.
[0028] In one possible implementation, after controlling the variable-focus optical lens according to the control strategy, the method further includes:
[0029] Collect the execution effect data fed back by the variable zoom optical lens, and optimize the control strategy and formulate the algorithm based on the execution effect data fed back by the variable zoom optical lens.
[0030] A third aspect of this application provides a visual control device for an intelligent learning machine, applied to a processor in a visual control system for an intelligent learning machine. The intelligent learning machine visual control system is an intelligent learning machine visual control system according to the first aspect or any implementation thereof, comprising:
[0031] The calibration parameter acquisition unit is used to calibrate the variable focal length optical lens using the calibration unit, so as to obtain the calibration parameters of the variable focal length optical lens corresponding to different learning scenarios of the intelligent learning machine;
[0032] The control unit is used to automatically identify the current learning scene and at least one observation target contained in the current learning scene during system operation; formulate a control strategy based on the identification result and control the variable focal length optical lens to observe the at least one observation target in sequence according to the control strategy; and directly apply the calibration parameters corresponding to the current learning scene to control the variable focal length optical lens during observation.
[0033] In one possible implementation, the calibration parameter acquisition unit is specifically used for:
[0034] Determine whether the preset calibration trigger conditions are met;
[0035] The variable-focus optical lens is controlled to optimize optical parameters for different distance calibration benchmarks, thereby determining the calibration parameters of the variable-focus optical lens for different learning scenarios of the intelligent learning machine.
[0036] In one possible implementation, the control unit is specifically used for:
[0037] The variable-focus optical lens is controlled to perform active scanning to acquire multiple frames of raw environmental images;
[0038] The multi-frame original environment images are analyzed to obtain the learning scene recognition result; the learning scene recognition result is used to indicate the current learning scene and at least one observation target contained in the current learning scene and the type corresponding to each observation target;
[0039] A control strategy is formulated based on the learning scene recognition results. The control strategy is used to indicate the observation plan and corresponding lens parameters for each target in the current environment.
[0040] The variable focal length optical lens is controlled according to the control strategy.
[0041] In one possible implementation, the device further includes: a control strategy formulation unit;
[0042] The control strategy formulation unit is used to evaluate the observation priority of each observation target according to the current learning scenario and the type corresponding to each observation target; generate the time sequence for observing each observation target and the corresponding lens parameters according to the observation priority; when the observation time windows of at least two observation targets overlap, conflict resolution is performed according to the observation priority to determine the actual observation time of each observation target.
[0043] In one possible implementation, the apparatus further includes: an optimization unit;
[0044] The optimization unit is used to collect the execution effect data fed back by the zoom optical lens after controlling the zoom optical lens according to the control strategy, and to optimize the control strategy and formulate an algorithm based on the execution effect data fed back by the zoom optical lens.
[0045] The fourth aspect of this application provides an intelligent learning machine, including the intelligent learning machine vision control system of the first aspect or any implementation thereof.
[0046] The fifth aspect of this application provides a computer program product, including computer-readable instructions, which, when executed on an intelligent learning machine, cause the intelligent learning machine to implement the intelligent learning machine visual control method of the second aspect or any implementation thereof.
[0047] The sixth aspect of this application provides a computer-readable storage medium carrying one or more computer programs that, when executed by an intelligent learning machine, enable the intelligent learning machine to implement the intelligent learning machine visual control method of the second aspect or any implementation thereof.
[0048] By employing the above technical solutions, the intelligent learning machine visual control system, method, and related devices provided in this application utilize a highly integrated, non-mechanical or minimally mechanically movable single variable-focus optical lens to replace the multiple cameras or single camera of a rotatable mechanical gimbal in traditional learning machines. A calibration unit ensures the physical performance of the single variable-focus optical lens throughout its entire lifecycle. A processor and scene recognition mechanism endow the single variable-focus optical lens with the ability to meet the diverse learning scenario task requirements of the intelligent learning machine. Therefore, this solution achieves high-performance interactive experience while reducing material costs through high hardware integration. Simultaneously, thanks to its software-level scene adaptation capabilities, it perfectly matches the complex and ever-changing usage scenarios of intelligent learning machines. Attached Figure Description
[0049] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0050] Figure 1 This application provides a schematic diagram of the structure of a vision control system for an intelligent learning machine.
[0051] Figure 2 A flowchart illustrating a visual control method for an intelligent learning machine provided in an embodiment of this application;
[0052] Figure 3 This application provides a flowchart illustrating a method for calibrating a variable-focus optical lens using the calibration unit to obtain calibration parameters of the variable-focus optical lens corresponding to different learning scenarios.
[0053] Figure 4 This is a flowchart illustrating a method for identifying the current learning scenario and directly applying calibration parameters corresponding to the current learning scenario to control the zoom optical lens during system operation, as provided in an embodiment of this application.
[0054] Figure 5 This is a schematic diagram of the structure of a visual control device for an intelligent learning machine provided in an embodiment of this application. Detailed Implementation
[0055] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0056] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0057] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0058] With the deep integration of "AI + Education", smart learning machines have become a core tool to assist K-12 students in their learning. Currently, most mainstream high-end learning machines are equipped with cameras to meet the visual perception needs of different learning scenarios, such as real-time monitoring of students' posture and attention to protect their eyesight and spinal health; and taking pictures of homework and test papers for intelligent diagnosis and correction.
[0059] Existing technologies mainly include two solutions. The first is a multi-camera fixed-function solution, where a wide-angle front-facing camera and a high-resolution rear or top-facing camera are independently installed on the device's frame. These two modules are completely independent in hardware and circuitry, and users manually switch between them via physical buttons or software buttons. The second is a rotatable mechanical gimbal single-camera solution, where a single camera is mounted on a motorized rotating gimbal. Users control the gimbal's rotation via commands, switching the camera's orientation between facing the user and the desktop. Essentially, this combines the functions of two fixed positions into a single camera through mechanical movement.
[0060] However, existing solutions still have significant drawbacks. Multi-camera solutions require the procurement, installation, and calibration of multiple camera modules, resulting in high material costs and assembly complexity, significantly increasing BOM costs. Multiple independent modules cannot operate simultaneously, lacking scene fusion capabilities. Multiple camera holes disrupt the device's integrity, affecting aesthetics and encroaching on internal space, hindering slim and lightweight designs. Rotatable mechanical gimbal solutions, on the other hand, suffer from wear and dust ingress risks in moving mechanical parts, are prone to failure with prolonged use, and are more vulnerable to drops. Gimbal rotation is accompanied by noise and delay (typically 1-2 seconds), preventing instantaneous and seamless scene switching, limiting dynamic performance, and causing long startup response times and image blurring during rotation.
[0061] Therefore, how to provide a highly integrated, cost-controllable, and adaptable visual control solution for intelligent learning machines that can adapt to different learning scenarios has become a technical problem that urgently needs to be solved by those skilled in the art.
[0062] To address the aforementioned problems, this application provides an intelligent learning machine vision control system. The intelligent learning machine vision control system of this application embodiment will be described in detail below with reference to the accompanying drawings.
[0063] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a visual control system for an intelligent learning machine provided in an embodiment of this application, as shown below. Figure 1 As shown in the embodiment of this application, a vision control system for an intelligent learning machine includes:
[0064] Monocular adaptive vision acquisition module and processor;
[0065] The monocular adaptive vision acquisition module includes a variable focal length optical lens and a calibration unit;
[0066] The processor is used to calibrate the variable-focus optical lens using the calibration unit to obtain calibration parameters corresponding to different learning scenarios of the intelligent learning machine; and during system operation, it automatically identifies the current learning scenario and at least one observation target contained in the current learning scenario; formulates a control strategy based on the identification result and controls the variable-focus optical lens to observe the at least one observation target in sequence according to the control strategy, and directly applies the calibration parameters corresponding to the current learning scenario to control the variable-focus optical lens during observation.
[0067] In one possible implementation, when the processor formulates a control strategy based on the recognition result, it is specifically used for:
[0068] Based on the current learning scenario and the type corresponding to each observation target, evaluate the observation priority of each observation target;
[0069] Based on the observation priority, the time sequence for observing each of the observation targets and the corresponding lens parameters are generated.
[0070] When the observation time windows of at least two of the observation targets overlap, conflict resolution is performed according to the observation priority to determine the actual observation time of each of the observation targets.
[0071] In one possible implementation, the processor is further configured to:
[0072] Collect the execution effect data fed back after the variable zoom optical lens performs observations according to the control strategy;
[0073] Based on the execution effect data, the algorithm for formulating the control strategy is optimized.
[0074] In this application, during system operation, the processor can control the variable zoom optical lens to switch lens parameters based on the calibration parameters corresponding to different learning scenarios.
[0075] The intelligent learning machine vision control system provided in this application replaces the multiple cameras or single camera of the rotating mechanical gimbal in traditional learning machines with a highly integrated, non-mechanical or minimally mechanically moving single zoom optical lens. The calibration unit ensures the physical performance of the single zoom optical lens throughout its entire life cycle. Through the processor and scene recognition mechanism, the single zoom optical lens is endowed with the ability to cope with the multi-learning scenario task requirements of the intelligent learning machine. Therefore, this solution achieves the reduction of material costs through high hardware integration while ensuring a high-performance interactive experience. At the same time, thanks to the scene adaptation capability at the software level, it perfectly matches the complex and ever-changing usage scenarios of the intelligent learning machine.
[0076] In one possible implementation, the variable-focus optical lens is a liquid lens.
[0077] In one possible implementation, the calibration unit has at least two built-in distance calibration benchmarks, with different distance calibration benchmarks corresponding to different learning scenarios and different calibration reference objects.
[0078] In one possible implementation, the calibration unit has a built-in dual-distance calibration reference, which includes a distance calibration reference for a first learning scenario and a distance calibration reference for a second learning scenario. The distance calibration reference for the first learning scenario corresponds to a first calibration reference, and the distance calibration reference for the second learning scenario corresponds to a second calibration reference.
[0079] In one possible implementation, a calibration reference can be created based on off-axis optical design, for example, by introducing a reference at a fixed distance into the optical path of the optical structure. This fixed distance can be set based on scenario requirements. Taking a learning machine's learning scenarios as an example, the first learning scenario corresponds to photographing the desktop, and the second learning scenario corresponds to photographing the user. The fixed distance of the distance calibration reference for the first learning scenario corresponds to a typical desktop working distance, and the fixed distance of the distance calibration reference for the second learning scenario corresponds to a typical user monitoring distance.
[0080] In one possible implementation, the intelligent learning machine vision control system further includes an environment interaction unit. In this case, the processor is also used to identify the current learning scene during system operation and, when directly applying the calibration parameters corresponding to the current learning scene to control the variable focal length optical lens, coordinate with the environment interaction unit to perform environmental adjustments.
[0081] In one possible implementation, when the processor uses the calibration unit to calibrate the zoom lens to obtain calibration parameters of the zoom lens corresponding to different learning scenarios, it is specifically used for:
[0082] Determine whether the preset calibration trigger conditions are met; control the variable-focus optical lens to optimize the optical parameters of the distance calibration reference respectively, and determine the calibration parameters of the variable-focus optical lens corresponding to different learning scenarios of the intelligent learning machine.
[0083] In this application, the preset calibration trigger condition is used to ensure that the zoom lens maintains accurate focusing performance under different environments and usage conditions. The specific implementation of the calibration trigger condition will be described in detail through the following embodiments.
[0084] In one possible implementation, when the processor identifies the current learning scene during system operation and directly applies the calibration parameters corresponding to the current learning scene to control the zoom optical lens, it is specifically used for:
[0085] The variable-focus optical lens is controlled to perform active scanning to acquire multiple frames of raw environmental images;
[0086] The multi-frame original environment images are analyzed to obtain the learning scene recognition result; the learning scene recognition result is used to indicate the targets included in the current environment and the learning scene corresponding to each target.
[0087] A control strategy is formulated based on the learning scene recognition results. The control strategy is used to indicate the observation plan and corresponding lens parameters of each target in the current environment. The zoom optical lens is controlled according to the control strategy.
[0088] In one possible implementation, the processor is further configured to collect execution effect data fed back by the zoom optical lens, and to optimize the control strategy and formulate an algorithm based on the execution effect data fed back by the zoom optical lens.
[0089] In this application, the execution effect data fed back by the variable zoom optical lens can be the image quality of the new image acquired after the variable zoom optical lens is adjusted.
[0090] Reference Figure 2 , Figure 2 This is a flowchart illustrating a visual control method for an intelligent learning machine provided in an embodiment of this application, as shown below. Figure 2 As shown in the figure, the execution subject of the intelligent learning machine vision control method provided in this application embodiment is the processor of the aforementioned intelligent learning machine vision control system. The method may include the following steps, which are described in detail below.
[0091] S201: The calibration unit is used to calibrate the variable-focus optical lens to obtain the calibration parameters of the variable-focus optical lens corresponding to different learning scenarios of the intelligent learning machine.
[0092] S202: During system operation, the current learning scene and at least one observation target contained in the current learning scene are automatically identified; a control strategy is formulated based on the identification result and the variable focal length optical lens is controlled to observe the at least one observation target in sequence according to the control strategy. During observation, the calibration parameters corresponding to the current learning scene are directly applied to control the variable focal length optical lens.
[0093] Reference Figure 3 , Figure 3 A flowchart illustrating a method for calibrating a zoomable optical lens using the calibration unit to obtain calibration parameters of the zoomable optical lens for different learning scenarios, as provided in this application embodiment, is shown below. Figure 3 As shown, the method may include the following steps, which are described in detail below.
[0094] S301: Determine whether the preset calibration trigger condition is met;
[0095] In this application, the preset calibration trigger condition is to ensure that the zoom lens maintains accurate focusing performance in different environments and usage conditions.
[0096] The liquid refractive index and driving voltage-curvature relationship of a zoom lens drift with temperature changes, making temperature variation a crucial calibration trigger factor. Therefore, in one possible implementation, preset calibration trigger conditions can be determined based on ambient temperature changes. For example, temperature changes exceeding a threshold, temperature crossing a specific range boundary, the first occurrence of extreme temperatures, and abnormal temperature change rates can all serve as preset calibration trigger conditions. For instance, recalibration is triggered when the ambient temperature changes by more than ±5°C compared to the last calibration. Preset key temperature nodes (e.g., 0°C, 25°C, 40°C) trigger calibration when the temperature crosses any of these nodes. A forced calibration is triggered when the module first experiences an environment below 0°C or above 45°C to establish baseline parameters under extreme conditions. Calibration is triggered when the temperature change rate per unit time exceeds a set value (e.g., 2°C / minute), which may indicate a drastic environmental change where existing parameters are no longer applicable.
[0097] Over time, zoom lenses may experience material fatigue, subtle changes in liquid properties, or minor drifts in their mechanical structure, requiring periodic recalibration. Therefore, in one possible implementation, preset calibration trigger conditions can be determined based on lens usage duration. For example, thresholds can be set for accumulated operating time, accumulated zoom counts, continuous operating time, and reaching a preset calibration cycle after waking from sleep mode. For instance, calibration could be triggered every 100 hours of cumulative power-on operation, every 5000 zoom operations, or more than 8 hours of continuous continuous operation. If more than 24 hours have passed since the last calibration after the module wakes from sleep mode, a quick calibration could be triggered.
[0098] In addition, to ensure the effectiveness of the zoom lens even when the device is in a "passive waiting" state, one possible implementation is that the preset calibration trigger conditions can be determined based on the most recent calibration time. For example, absolute time interval thresholds, combined calendar time and usage time triggers, mandatory calibration before critical events, and periodic health checks can all be used as preset calibration trigger conditions. For instance, if more than 7 days (168 hours) have passed since the last calibration, a calibration is forcibly triggered regardless of other conditions. This is a fallback mechanism to prevent parameter drift caused by prolonged inactivity. If the calendar time exceeds 3 days and the cumulative usage time exceeds 10 hours, calibration is triggered. Before the device is powered on or enters high-precision mode (such as enabling fingertip word lookup or job correction functions), if more than 2 hours have passed since the last calibration, a calibration is triggered to ensure accuracy in critical scenarios. The module has a built-in timer that automatically triggers a lightweight calibration once every morning (when the device is idle) for daily calibration and health monitoring.
[0099] S302: Control the variable-focus optical lens to optimize the optical parameters of different distance calibration benchmarks, and determine the calibration parameters of the variable-focus optical lens corresponding to different learning scenarios of the intelligent learning machine.
[0100] In this application, the specific process of controlling the variable-focus optical lens to optimize the optical parameters of a distance calibration benchmark and determining the calibration parameters of the variable-focus optical lens corresponding to the learning scene of the distance calibration benchmark can be as follows: control the variable-focus optical lens to align with the calibration reference object corresponding to the distance calibration benchmark and apply voltage scanning, evaluate the image sharpness by calculating the image gradient or contrast, and record the voltage value corresponding to the sharpness peak as the calibration parameters of the variable-focus optical lens corresponding to the learning scene of the distance calibration benchmark.
[0101] For ease of understanding, assume that the calibration unit has a built-in dual-distance calibration reference, which includes a distance calibration reference for a first learning scene and a distance calibration reference for a second learning scene. The distance calibration reference for the first learning scene corresponds to a first calibration reference, and the distance calibration reference for the second learning scene corresponds to a second calibration reference. Specifically, controlling the variable-focus optical lens to optimize the optical parameters of both the first and second learning scene distance calibration references to determine the calibration parameters corresponding to different learning scenes can be achieved by controlling the variable-focus optical lens to optimize the optical parameters of both the first and second learning scene distance calibration references respectively.
[0102] The process of controlling the variable-focus optical lens to optimize the optical parameters of the distance calibration references of the first learning scene and the second learning scene can be as follows: control the variable-focus optical lens to align with the first calibration reference and apply voltage scanning, evaluate the image sharpness by calculating the image gradient or contrast, record the first voltage value corresponding to the sharpness peak as the calibration parameter of the variable-focus optical lens corresponding to the first learning scene, then control the variable-focus optical lens to align with the second calibration reference and apply voltage scanning, evaluate the image sharpness by calculating the image gradient or contrast, and record the second voltage value corresponding to the sharpness peak as the calibration parameter of the variable-focus optical lens corresponding to the second learning scene.
[0103] Based on the above calibration scheme, a long-term stable optical reference can be established for the zoom lens, overcoming parameter drift caused by ambient temperature and device aging.
[0104] Reference Figure 4 , Figure 4 This application provides a flowchart illustrating a method for automatically identifying the current learning scene and at least one observation target included in the current learning scene during system operation; formulating a control strategy based on the identification result and controlling the variable-focus optical lens to sequentially observe the at least one observation target according to the control strategy; and directly applying calibration parameters corresponding to the current learning scene to control the variable-focus optical lens during observation. Figure 4 As shown, the method may include the following steps, which are described in detail below.
[0105] S401: Control the variable zoom optical lens to perform active scanning to acquire multiple frames of raw environmental images;
[0106] In this application, a layered scanning strategy can be used to control the variable-focus optical lens to actively scan and acquire multiple frames of original environmental images. Specifically, on the one hand, the variable-focus optical lens can be controlled to use a wide-angle setting to quickly acquire an overview of the environment, identify the distribution of people and major objects, and achieve the purpose of rapid global perception. On the other hand, the variable-focus optical lens can be controlled to perform zoom scanning on the identified region of interest to acquire detailed feature information and achieve the purpose of target detail recognition.
[0107] S402: Analyze the multiple frames of original environment images to obtain the learning scene recognition result; the learning scene recognition result is used to indicate the current learning scene and at least one observation target contained in the current learning scene and the type corresponding to each observation target;
[0108] In this application, environmental feature extraction, target recognition and classification, and target spatial relationship modeling can be performed on multiple frames of original environmental images to obtain learning scene recognition results. Environmental feature extraction can be the extraction of key visual features from multiple frames of original environmental images. Target recognition and classification can be the identification and classification of key elements (such as users, books, notebooks, mobile phones, stationery, etc.) in the learning scene based on key visual features. Target spatial relationship modeling can be the establishment of spatial positional relationships and relative distances between key elements.
[0109] S403: Formulate a control strategy based on the learning scene recognition results. The control strategy is used to indicate the observation plan and corresponding lens parameters for each target in the current environment.
[0110] In this application, a control strategy formulation algorithm can be preset, which formulates a control strategy based on the learning scene recognition results. In one possible implementation, the observation priority of each observation target can be evaluated based on the current learning scene and the type corresponding to each observation target; based on the observation priority, the time sequence of observation for each observation target and the corresponding lens parameters are generated; when the observation time windows of at least two observation targets overlap, conflict resolution is performed based on the observation priority to determine the actual observation time of each observation target.
[0111] Specifically, the control strategy formulation algorithm performs the following steps:
[0112] Step 1, Multi-target recognition and classification: Extract all observed targets in the current environment from the learning scene recognition results, and classify each observed target.
[0113] For example, when the recognition results are "user face", "book page", "finger tip", and "workbook", they are classified as "user monitoring target", "desktop content target" and "operation interaction target", respectively.
[0114] Step 2, Observation Priority Assessment: Based on the current learning scenario type and the category attributes of each observation target, assess the observation priority of each observation target.
[0115] For example, when the current learning scenario type is "finger-tip reading," the "finger tip" and its adjacent "book page" text area in the operation interaction category are given the highest priority; when the current learning scenario type is "posture detection," the "user's face" is given the highest priority. The priority evaluation also considers user state factors. For example, when the user's hand is detected moving quickly toward the book, the priority of the "finger tip" is dynamically increased to predict the upcoming "finger-tip reading" behavior.
[0116] Step 3, Control Strategy Generation: Based on the observation priority of each observation target, dynamically plan the observation sequence and corresponding lens parameters for each observation target to generate a control strategy.
[0117] The control strategy includes a time-ordered sequence of lens control instructions, each containing a target identifier, a target observation time period, and lens parameters for that time period. For example, if the current priority order is "finger tip (highest)" > "text area on book page (second highest)" > "user face (normal)", then the control strategy can be planned as follows: from 0-50ms, switch to macro calibration parameters to capture a high-resolution image of the fingertip; from 50-120ms, switch to desktop document calibration parameters to capture an image of the book page; and from 120-180ms, switch to user monitoring calibration parameters to capture an image of the user's face.
[0118] Step 4, Conflict Resolution and Strategy Optimization: When there are multiple observation targets and the observation time windows overlap, conflict resolution is carried out according to the observation priority to determine the actual observation time of each observation target.
[0119] For example, if the observation time windows of "finger tip" and "text area of book page" overlap, but cannot be collected simultaneously due to differences in lens parameters, the processor decides to execute the observation of "finger tip" first and then the observation of "text area of book page" according to priority, and records the conflict situation for subsequent strategy optimization.
[0120] In one possible implementation, the control strategy formulation algorithm also supports a trial strategy in low-confidence scenarios: when the confidence of the scene recognition result is lower than a preset threshold, the processor cannot determine the unique main scene type. At this time, a multi-candidate parameter trial strategy is adopted—controlling the variable focal length optical lens to try multiple candidate calibration parameters in sequence for rapid acquisition. By comparing the sharpness, target coverage and other indicators of the acquired image, the optimal parameters are determined before subsequent continuous observation is performed.
[0121] Based on the above control strategy formulation method, this application provides several types of learning scenarios for intelligent learning machines, as well as the visual tasks, observation targets, calibration parameters, and control strategies corresponding to different learning scenarios, as follows:
[0122] Scenario 1: Posture and Attention Monitoring
[0123] The visual task in this scenario is to monitor the user's head position, body tilt, and eye opening and closing status in real time to determine whether the user is leaning on the table, leaving the seat, or having their attention diverted.
[0124] The observation targets in this scenario are: the user's face and upper body.
[0125] The typical object distance for this scenario is 40cm to 70cm.
[0126] The calibration parameters corresponding to this scenario are: user monitoring calibration parameters.
[0127] The control strategy in this scenario is as follows: control the zoom lens to maintain the mid-range focal length, quickly autofocus, and continuously track the user's head; when a significant displacement of the user's head position is detected, trigger refocusing.
[0128] Scenario 2: Finger-tip reading and word lookup scenario
[0129] The visual task in this scenario is to accurately identify the text at the location pointed to by the user's finger and translate or interpret it in real time.
[0130] The observation targets in this scene are: the fingertips and the text on the book below the fingertips.
[0131] The typical object distance for this scene is 15cm to 30cm.
[0132] The calibration parameters corresponding to this scenario are: desktop macro calibration parameters.
[0133] The control strategy in this scenario is as follows: control the zoom lens to quickly zoom to macro mode; if the zoom lens supports a variable aperture, then reduce the aperture to increase the depth of field, ensuring that the fingertip and adjacent text are both clear at the same time.
[0134] Scenario 3: Intelligent Grading of Homework and Test Papers
[0135] The visual task in this scenario is to clearly image the workbooks or test papers laid out on the table for subsequent OCR recognition and intelligent grading.
[0136] The observation target in this scenario is a full-page document, such as an A4-sized page.
[0137] The typical object distance for this scenario is 30cm to 50cm.
[0138] The calibration parameters for this scenario are: desktop document calibration parameters.
[0139] The lens control strategy in this scenario is to adjust the focal length of the zoom lens to bring the entire document into the frame, and optimize the focus plane to ensure that the text in the center and at the edges of the image is clear at the same time.
[0140] Scenario 4: Environment-Adaptive Data Acquisition Scenario
[0141] The visual task in this scenario is to ensure image quality under complex lighting conditions such as low light and backlight.
[0142] The observation target in this scenario is the main target in the current learning scenario. For example, when reading with fingertips in low light, the main target is the fingertip and the text in the book.
[0143] The typical object distance for this scenario is: depending on the main scene. For example, if the main scene is fingertip reading, the object distance is 15cm to 30cm, and if the main scene is posture monitoring, the object distance is 40cm to 70cm.
[0144] The calibration parameters corresponding to this scenario are: the calibration parameters corresponding to the current learning scenario. For example, when the current learning scenario is fingertip reading, the desktop macro calibration parameters are used, and when the current learning scenario is posture monitoring, the user monitoring calibration parameters are used.
[0145] The control strategy in this scenario is as follows: the environmental interaction unit is coordinated to adjust the environment, including but not limited to the fill light, screen brightness and screen color temperature; after adjusting the lighting conditions to meet the preset requirements, the zoom lens is then controlled to perform focusing and subsequent observation.
[0146] Scenario 5: Multi-objective hybrid scenario
[0147] The visual task in this scenario is to simultaneously handle multiple visual tasks, such as monitoring posture while waiting for fingertip reading instructions.
[0148] The observation target in this scenario is multiple targets, such as a user, a book, and a finger, all existing simultaneously.
[0149] The typical object distance in this scenario is as follows: different objects have different distances, for example, the distance to a user's face is 40cm to 70cm, the distance to a book is 30cm to 50cm, and the distance to a fingertip is 15cm to 30cm.
[0150] The calibration parameters for this scenario are: multiple calibration parameters are rotated, that is, user monitoring calibration parameters, desktop document calibration parameters and desktop macro calibration parameters are used at different time periods.
[0151] The control strategies in this scenario are as follows: a time-slice rotation strategy is adopted, which periodically switches between multiple calibration parameters according to the order of observation priority; or an event-triggered strategy is adopted, which prioritizes switching to the calibration parameters corresponding to a target when a state change is detected, and suspends the observation of other targets until the current task is completed.
[0152] It should be noted that the learning scenarios of the intelligent learning machine are not limited to the above-mentioned ones. Other learning scenarios and control strategies generated based on the above-mentioned control strategy generation logic are all within the protection scope of this invention.
[0153] S404: Control the zoom optical lens according to the control strategy.
[0154] In this application, the zoomable optical lens can be controlled to switch lens parameters at appropriate times to observe various targets in the current environment. It should be noted that during the process of controlling the zoomable optical lens according to the control strategy, abnormal situations can also be detected. When an abnormal situation is detected, the control strategy is dynamically adjusted, and the zoomable optical lens is controlled based on the adjusted control strategy. In this application, abnormal situations include, but are not limited to, user departure, sudden changes in light, etc., and this application does not impose any limitations on these.
[0155] In one possible implementation, after controlling the variable-focus optical lens according to the control strategy, the method further includes:
[0156] Collect the execution effect data fed back by the variable zoom optical lens, and optimize the control strategy and formulate the algorithm based on the execution effect data fed back by the variable zoom optical lens.
[0157] In this application, the execution effect data fed back by the variable zoom optical lens can be the image quality of the new image acquired after the variable zoom optical lens is adjusted.
[0158] The core of this application lies in constructing a closed-loop "perception-decision-execution" system specifically designed for intelligent learning machine learning scenarios. By forming a closed-loop data flow through raw image data, scene recognition results, and control strategies, the system is able to proactively perceive and analyze complex learning environments. It can identify the spatial relationships and semantic information of multiple targets, achieving a deep coupling between the stability of optical accuracy and the intelligence of learning scenario understanding, and realizing intelligent visual resource scheduling in multiple learning scenarios.
[0159] The above describes a visual control method for an intelligent learning machine provided by an embodiment of this application. The following describes an apparatus for executing the above-described visual control method for an intelligent learning machine. This apparatus is applied to a processor in an intelligent learning machine visual control system, which is any of the aforementioned intelligent learning machine visual control systems.
[0160] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a visual control device for an intelligent learning machine provided in an embodiment of this application. Figure 5 As shown, the visual control device for the intelligent learning machine includes:
[0161] The calibration parameter acquisition unit 11 is used to calibrate the variable focal length optical lens using the calibration unit, so as to obtain the calibration parameters of the variable focal length optical lens corresponding to different learning scenarios of the intelligent learning machine.
[0162] The control unit 12 is used to automatically identify the current learning scene and at least one observation target contained in the current learning scene during system operation; formulate a control strategy based on the identification result and control the variable focal length optical lens to observe the at least one observation target in sequence according to the control strategy; and directly apply the calibration parameters corresponding to the current learning scene to control the variable focal length optical lens during observation.
[0163] In one possible implementation, the calibration parameter acquisition unit is specifically used for:
[0164] Determine whether the preset calibration trigger conditions are met;
[0165] The variable-focus optical lens is controlled to optimize optical parameters for different distance calibration benchmarks, thereby determining the calibration parameters of the variable-focus optical lens for different learning scenarios of the intelligent learning machine.
[0166] In one possible implementation, the control unit is specifically used for:
[0167] The variable-focus optical lens is controlled to perform active scanning to acquire multiple frames of raw environmental images;
[0168] The multi-frame original environment images are analyzed to obtain the learning scene recognition result; the learning scene recognition result is used to indicate the targets included in the current environment and the learning scene corresponding to each target.
[0169] A control strategy is formulated based on the learning scene recognition result. The control strategy is used to indicate the current learning scene, at least one observation target contained in the current learning scene, and the type corresponding to each observation target.
[0170] The variable focal length optical lens is controlled according to the control strategy.
[0171] In one possible implementation, the device further includes: a control strategy formulation unit;
[0172] The control strategy formulation unit is used to evaluate the observation priority of each observation target according to the current learning scenario and the type corresponding to each observation target; generate the time sequence for observing each observation target and the corresponding lens parameters according to the observation priority; when the observation time windows of at least two observation targets overlap, conflict resolution is performed according to the observation priority to determine the actual observation time of each observation target.
[0173] In one possible implementation, the apparatus further includes: an optimization unit;
[0174] The optimization unit is used to collect the execution effect data fed back by the zoom optical lens after controlling the zoom optical lens according to the control strategy, and to optimize the control strategy and formulate an algorithm based on the execution effect data fed back by the zoom optical lens.
[0175] Each unit in the aforementioned intelligent learning machine visual control device can be implemented entirely or partially through software, hardware, or a combination thereof. These units can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each unit.
[0176] This application also provides an intelligent learning machine, including any of the intelligent learning machine vision control systems provided in this application.
[0177] This application also provides a computer program product, including computer-readable instructions, which, when executed on a smart learning machine, enable the smart learning machine to implement any of the smart learning machine visual control methods provided in this application.
[0178] This application also provides a computer-readable storage medium carrying one or more computer programs. When the one or more computer programs are executed by an intelligent learning machine, the intelligent learning machine can implement any of the intelligent learning machine visual control methods provided in this application.
[0179] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0180] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0181] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0182] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A visual control system for an intelligent learning machine, characterized in that, include: Monocular adaptive vision acquisition module and processor; The monocular adaptive vision acquisition module includes a variable focal length optical lens and a calibration unit; The processor is used to calibrate the variable-focus optical lens using the calibration unit to obtain calibration parameters of the variable-focus optical lens corresponding to different learning scenarios of the intelligent learning machine; During system operation, the system automatically identifies the current learning scene and at least one observation target contained in the current learning scene; based on the identification results, a control strategy is formulated and the variable focal length optical lens is controlled to observe the at least one observation target in sequence according to the control strategy. During observation, the calibration parameters corresponding to the current learning scene are directly applied to control the variable focal length optical lens.
2. The system according to claim 1, characterized in that, The zoomable optical lens is a liquid lens.
3. The system according to claim 1, characterized in that, The calibration unit has at least two built-in distance calibration benchmarks, and different distance calibration benchmarks correspond to different learning scenarios and different calibration reference objects.
4. The intelligent learning machine vision control system according to claim 1, characterized in that, The system also includes an environment interaction unit. In this case, the processor is also used to identify the current learning scene during system operation and, when directly applying the calibration parameters corresponding to the current learning scene to control the variable focal length optical lens, coordinate with the environment interaction unit to perform environmental adjustments.
5. A visual control method for an intelligent learning machine, characterized in that, A processor applied in a vision control system for an intelligent learning machine, wherein the intelligent learning machine vision control system is the intelligent learning machine vision control system according to any one of claims 1 to 4, comprising: The variable-focus optical lens is calibrated using a calibration unit to obtain calibration parameters for the variable-focus optical lens and different learning scenarios of the intelligent learning machine. During system operation, the system automatically identifies the current learning scene and at least one observation target contained in the current learning scene; based on the identification results, a control strategy is formulated and the variable focal length optical lens is controlled to observe the at least one observation target in sequence according to the control strategy. During observation, the calibration parameters corresponding to the current learning scene are directly applied to control the variable focal length optical lens.
6. The method according to claim 5, characterized in that, The calibration of the variable-focus optical lens using the calibration unit to obtain calibration parameters for the variable-focus optical lens corresponding to different learning scenarios includes: Determine whether the preset calibration trigger conditions are met; The variable-focus optical lens is controlled to optimize optical parameters for different distance calibration benchmarks, thereby determining the calibration parameters of the variable-focus optical lens for different learning scenarios of the intelligent learning machine.
7. The method according to claim 5, characterized in that, During system operation, the system automatically identifies the current learning scene and at least one observation target contained within it; based on the identification results, a control strategy is formulated, and the variable-focus optical lens is controlled to sequentially observe the at least one observation target according to the control strategy. During observation, calibration parameters corresponding to the current learning scene are directly applied to control the variable-focus optical lens, including: The variable-focus optical lens is controlled to perform active scanning to acquire multiple frames of raw environmental images; The multi-frame original environmental images are analyzed to obtain the learning scene recognition result; the learning scene recognition result is used to indicate the current learning scene and at least one observation target contained in the current learning scene and the type corresponding to each observation target; A control strategy is formulated based on the learning scene recognition results. The control strategy is used to indicate the observation plan and corresponding lens parameters for each target in the current environment. The variable focal length optical lens is controlled according to the control strategy.
8. The method according to claim 7, characterized in that, The step of formulating a control strategy based on the learning scene recognition results includes: Based on the current learning scenario and the type corresponding to each observation target, evaluate the observation priority of each observation target; Based on the observation priority, the time sequence for observing each of the observation targets and the corresponding lens parameters are generated. When the observation time windows of at least two of the observation targets overlap, conflict resolution is performed according to the observation priority to determine the actual observation time of each of the observation targets.
9. The method according to claim 7 or 8, characterized in that, After controlling the variable-focus optical lens according to the control strategy, the method further includes: Collect the execution effect data fed back by the variable zoom optical lens, and optimize the control strategy and formulate the algorithm based on the execution effect data fed back by the variable zoom optical lens.
10. An intelligent learning machine, characterized in that, Includes the intelligent learning machine vision control system as described in any one of claims 1 to 4.