Intelligent control system for AI interactive exhibition study exhibition stand
By using weight sensors and cameras to collaboratively perceive user group attributes, and combining AI technology with environmental interference calibration, the exhibition booth achieves intelligent interaction, solving the problems of inaccurate user identification and inaccurate judgment of environmental interference in existing technologies, and improving the interactive reliability and user experience of the booth.
Patent Information
- Application Number
- CN202511471856.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-02-03
AI Technical Summary
Existing interactive exhibition booths struggle to accurately identify user group attributes and adapt content modes in complex environments, resulting in low engagement and a lack of environmental interference calibration mechanisms, which negatively impacts the interactive experience.
By using weight sensors and cameras to collaboratively perceive users, AI identifies user group attributes, dynamically adjusts gesture and voice recognition thresholds, and combines environmental interference assessment parameters to construct preset marked areas and divide them into sub-areas for environmental interference calibration, thereby achieving intelligent control of multimedia devices.
It accurately identifies user groups and loads adapted content, reduces environmental interference and misjudgment, improves interaction reliability and user experience, supports flexible replacement of booth content and format, and adapts to complex exhibition scenarios.
Smart Images

Figure CN121462736A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of intelligent control, in particular to an AI interactive exhibition learning kiosk intelligent control system. BACKGROUND
[0002] In the current digital upgrading process of an exhibition hall, in order to quickly change the exhibition content and form under the condition that the original kiosk remains unchanged (saving time and cost), the natural center designs and creates a multifunctional AI interactive exhibition learning kiosk. The interactive exhibition learning kiosk has become an important carrier for popular science communication, but the prior art is still difficult to adapt to complex actual application scenarios.
[0003] For example, a physics and mechanics popular science kiosk of a science and technology museum adopts infrared induction triggering interaction, is matched with fixed text and video exhibition learning content, and gesture and voice recognition thresholds are factory default values. When 3 to 5 children groups approach on weekends, the infrared induction is frequently triggered to start due to the obstruction of surrounding tourists, and the loaded content is formula derivation content for adult groups, so that children cannot understand, leading to low interactive willingness. Meanwhile, under the environmental interference of the crowd noise in the exhibition hall and the direct sunlight on the kiosk screen in the afternoon, the preset low gesture recognition threshold makes children randomly waving hands be misjudged as content switching instructions, and the voice recognition cannot respond to the child voice instructions for explaining the principle of levers due to the fixed high threshold, which seriously affects the interactive experience. This case exposes the core defects of the prior art, that is, the lack of precise user perception and group attribute recognition ability of multiple sensors, the inability to dynamically adapt the content mode, and the lack of interference calibration mechanism based on real-time environmental characteristics. The fixed recognition threshold is difficult to balance the interactive accuracy and sensitivity in different environments. SUMMARY
[0004] The technical problem to be solved by the application is to provide an AI interactive exhibition learning kiosk intelligent control system, which realizes exhibition persistent application and flexible replacement of kiosk content and form, and adapts to exhibition learning needs in science and technology museum and museum scenes.
[0005] The multifunctional kiosk is a special device specially designed for product replacement. The characteristics are easy to change and replace, and the replaceable device combination design forms the advantage of depression exhibition replacement, including mutually replaceable and magic mode exhibition changes of the same size of multimedia, light boxes, model groups, micro scenes, exhibition cabinets and the like, so that new exhibitions appear in new forms in content and form under the condition that the kiosk remains unchanged.
[0006] To solve the above technical problems, the technical scheme of the application is as follows: In a first aspect, an AI interactive exhibition learning kiosk intelligent control method, the method comprising: According to the weight sensor, the user arrival is perceived and the process is started, the user group attribute is identified through the camera, and the adaptive exhibition learning content mode is loaded. After the content mode is loaded, environmental visual and audio information is collected to obtain initial environmental interference evaluation parameters; the camera takes three fixed physical structures on the exhibition stand as monitoring points, and the three monitoring points are the upper left corner of the interactive main screen, the upper right corner of the control panel, and the central marker point of the front edge of the exhibition stand table, to obtain a preset marker area; The preset marker area is divided into multiple regular sub-areas; the imaging feature changes of each sub-area in the visual picture are analyzed to obtain an environmental adjustment coefficient; the initial environmental interference evaluation parameters are calibrated by the environmental adjustment coefficient to obtain calibrated environmental interference evaluation parameters; According to the calibrated environmental interference evaluation parameters, the confidence threshold of gesture trajectory and semantic understanding is dynamically adjusted to obtain an adjusted gesture trajectory confidence threshold and a semantic understanding confidence threshold; Based on the adjusted gesture trajectory confidence threshold and the semantic understanding confidence threshold, a user gesture or voice instruction is recognized; the voice instruction confidence exceeding the threshold is effective; the gesture instruction needs to simultaneously satisfy the confidence threshold and pass the verification of the intersection of the gesture trajectory and the preset line segment to confirm the effective instruction, thereby controlling the multimedia device to display popular science content; After the popular science content display is completed, a user interaction behavior data is analyzed to obtain an attention index; according to the attention index, the loading priority of the exhibition content is optimized next time.
[0007] Further, according to the weight sensor sensing the arrival of a user and starting the process, the user group attribute is identified through the camera, and the adapted exhibition content mode is loaded, including: A trigger signal sent by a weight sensor arranged on the exhibition stand table is received, and the trigger signal is used to indicate that a user arrives at the exhibition stand interaction area; After receiving the trigger signal, a start instruction is sent to the camera of the exhibition stand, and user image data collected by the camera is received; The user image data is analyzed for group attribute recognition, and the group attribute at least includes an age level and a group size to obtain the recognized group attribute; According to the recognized group attribute, an exhibition content mode adapted thereto is matched and loaded from a preset standardized exhibition content library; based on AI control technology, a built-in question and answer engine is called and combined with visual AI analysis results to optimize the display content, to drive the multimedia device and the standardized exhibition to dynamically switch and control, to realize intelligent interaction without traditional teacher participation and adapt to changes in exhibition theme.
[0008] Further, after the content mode is loaded, the environmental visual and audio information is collected to obtain initial environmental interference evaluation parameters; the camera takes three fixed physical structures on the exhibition stand as monitoring points, the three monitoring points are the upper left corner of the interactive main screen, the upper right corner of the control panel and the central marker point of the front edge of the exhibition stand table top, to obtain a preset marker area, including: After the loading of the exhibition content mode is completed, the camera and the audio collection device are instructed to collect the visual picture and audio information of the current environment; based on the intensity and characteristics of the collected information, fusion calculation is performed to obtain initial environmental interference evaluation parameters; Based on the initial environmental interference evaluation parameters, environmental interference analysis is performed, and pre-stored exhibition stand physical structure coordinate data is called, which accurately corresponds to the preset positions of the upper left corner of the interactive main screen, the upper right corner of the control panel and the central marker point of the front edge of the exhibition stand table top in the camera picture; Based on the obtained preset position coordinates of the three monitoring points, the coordinates are taken as vertices to construct a triangular area covering the three monitoring points through geometric operation, and finally the triangular area is defined as the preset marker area.
[0009] Further, the preset marker area is divided into a plurality of regular sub-areas; the imaging feature changes of each sub-area in the visual picture are analyzed to obtain an environmental adjustment coefficient; the initial environmental interference evaluation parameters are calibrated through the environmental adjustment coefficient to obtain calibrated environmental interference evaluation parameters, including: After obtaining the preset marker area, the triangular area is divided into a plurality of rectangular sub-areas with equal areas according to the predefined grid rules; Based on the divided sub-areas, the imaging features of each sub-area in the visual picture in the continuous time sequence are extracted in turn, the imaging features include average brightness, color distribution and pixel motion vector, and the change amount of each sub-area imaging feature relative to the initial reference frame is calculated to obtain the statistical result of the imaging feature change amount; Based on the statistical result of the imaging feature change amount of all sub-areas, a comprehensive environmental adjustment coefficient is calculated through weighted fusion; the environmental adjustment coefficient is multiplied by the obtained initial environmental interference evaluation parameters to obtain the calibrated environmental interference evaluation parameters.
[0010] Further, according to the calibrated environmental interference evaluation parameters, the confidence threshold of gesture trajectory and semantic understanding is dynamically adjusted to obtain adjusted gesture trajectory confidence threshold and semantic understanding confidence threshold, including: Based on the calibrated environmental interference evaluation parameters, a pre-defined mapping relationship table is queried; the mapping relationship table defines the corresponding relationship between the environmental interference evaluation parameter value and the confidence threshold adjustment amount, so as to obtain the basic adjustment amount of the gesture trajectory confidence and the basic adjustment amount of the semantic understanding confidence respectively; The base adjustment amount of the gesture trajectory confidence is superimposed with a preset gesture trajectory confidence base threshold to obtain an adjusted gesture trajectory confidence threshold; The base adjustment amount of the semantic understanding confidence is superimposed with a preset semantic understanding confidence base threshold to obtain an adjusted semantic understanding confidence threshold.
[0011] Further, based on the adjusted gesture trajectory confidence threshold and the semantic understanding confidence threshold, the user gesture or voice instruction is identified; the voice instruction confidence exceeding the threshold is effective; the gesture instruction needs to simultaneously satisfy the confidence threshold and pass the verification of the intersection judgment of the gesture trajectory and the preset line segment to confirm the effective instruction, thereby controlling the multimedia device to display the popular science content, including: Based on the obtained adjusted gesture trajectory confidence threshold and the semantic understanding confidence threshold, the gesture of the real-time video stream collected by the camera and the voice of the real-time audio stream collected by the audio collection device are identified and analyzed to obtain real-time gesture recognition results and voice recognition analysis results; When the confidence of a certain voice instruction in the voice recognition analysis result exceeds the adjusted semantic understanding confidence threshold, the voice instruction is determined to be an effective instruction; The confidence of the gesture trajectory in the gesture recognition result is judged whether it exceeds the adjusted gesture trajectory confidence threshold; if it exceeds, the motion trajectory of the gesture is further subjected to geometric intersection judgment with a virtual line segment preset in the interaction space to obtain a geometric intersection judgment result; According to the geometric intersection judgment result, only when the motion trajectory of a certain gesture instruction intersects with the preset virtual line segment, the gesture instruction is finally determined to be an effective instruction; After determining any effective instruction, a control signal corresponding to the effective instruction is generated and sent to the multimedia device of the exhibition stand; wherein the control signal is used to drive the multimedia device to display the popular science content corresponding to the instruction, and according to the current exhibition theme requirement, the extension and lifting of the display screen and the opening and closing of the interactive multifunctional drawer are dynamically adjusted, and the exhibition equipment is combined to realize optimized space utilization and audience interaction effect.
[0012] Further, after the popular science content display is completed, the attention index is obtained by analyzing the user interaction behavior data; according to the attention index, the loading priority of the exhibition content is optimized next time, including: After the popular science content display is completed, the user interaction behavior data in the current interaction process is collected; Based on the collected user interaction behavior data, the attention index for quantifying the user's interest in different popular science content is obtained by comprehensive calculation through a weight calculation formula; According to the calculated attention index, the loading priority weight of the corresponding science popularization content in the content library is dynamically updated to obtain an updated loading priority weight. Based on the updated loading priority weight, in combination with the display device adopting a standardized interface, the optimized content loading sequence is persistently stored to a database; wherein the standardized display device supports flexible replacement of display content and physical form, so that the exhibition stand adapts to different exhibition themes according to the updated priority, and realizes persistent application of the exhibition.
[0013] In a second aspect, an AI interactive exhibition learning exhibition stand intelligent control system comprises: An acquisition module is configured to perceive the arrival of a user according to a weight sensor and start a process, identify user group attributes through a camera, and load an adapted exhibition learning content mode. An adjustment module is configured to, after loading the content mode, collect environmental visual and audio information to obtain initial environmental interference evaluation parameters; the camera takes three fixed physical structures on the exhibition stand as monitoring points, the three monitoring points are a top-left corner of an interactive main screen, a top-right corner of a control panel, and a central marker point at a front edge of a stand surface, to obtain a preset marker area; the preset marker area is divided into multiple regular sub-areas; imaging feature changes of each sub-area in a visual picture are analyzed to obtain an environmental adjustment coefficient; the initial environmental interference evaluation parameters are calibrated through the environmental adjustment coefficient to obtain calibrated environmental interference evaluation parameters. A calculation module is configured to, according to the calibrated environmental interference evaluation parameters, dynamically adjust a confidence threshold of gesture trajectory and semantic understanding to obtain an adjusted gesture trajectory confidence threshold and a semantic understanding confidence threshold; based on the adjusted gesture trajectory confidence threshold and the semantic understanding confidence threshold, a user gesture or voice instruction is recognized; a voice instruction confidence exceeding a threshold is effective; a gesture instruction needs to simultaneously satisfy a confidence threshold and pass a verification of intersection of a gesture trajectory and a preset line segment to confirm an effective instruction, thereby controlling a multimedia device to display science popularization content. A processing module is configured to, after the science popularization content display ends, analyze user interaction behavior data to obtain an attention index; and according to the attention index, optimize a loading priority of exhibition learning content at the next start.
[0014] In a third aspect, a computing device comprises: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method.
[0015] In a fourth aspect, a computer readable storage medium stores a program, when the program is executed by a processor, the method is implemented.
[0016] The above scheme of the present application at least includes the following beneficial effects: Through the cooperative perception of the weight sensor and the camera, the problem of false triggering of traditional infrared sensing easily affected by shielding is accurately avoided, the user age level and group size are accurately identified, the AI control technology, the built-in question and answer engine and the visual AI analysis module are combined, the adaptive exhibition content mode is loaded and the content is displayed through algorithm optimization, the intelligent interaction between the exhibition stand and the user can be realized without the participation of traditional teachers, the problem of low interaction willingness caused by the homogenization of existing exhibition stand content and the mismatch with user cognitive needs is solved, dynamic switching and control of multimedia, models and other exhibition devices are supported, the display screen can also be dynamically adjusted in terms of stretching, lifting, interactive multifunctional drawer opening and closing, forming a changeable, heuristic and multifunctional easy-to-operate exhibition stand combination, realizing a comprehensive AI interactive exhibition teaching aid, optimizing space utilization and audience interaction effect; at the same time, a preset marking area is constructed by taking the fixed physical structure of the exhibition stand as a monitoring point, an environmental adjustment coefficient is generated by analyzing the imaging characteristics of the sub-area, the dynamic calibration of the initial environmental interference evaluation parameter is realized, the accuracy of environmental interference judgment is improved, the confidence threshold of gesture trajectory and semantic understanding is dynamically adjusted according to the mapping relationship table of the calibrated parameters, which not only reduces the risk of misjudgment in high-interference environments, but also avoids the problem of missed judgment in low-interference environments, and the double verification mechanism of the intersection of gesture trajectory and preset line segment further guarantees the accuracy of command recognition; in addition, the attention index is generated by quantitatively analyzing user interaction behavior data, the exhibition content loading priority is dynamically optimized and is stored persistently, the high-attention content can be loaded preferentially when the same group attribute user is recognized next time, forming a personalized recommendation closed loop from perception to interaction to optimization, and the system establishes standardized exhibition devices to meet the needs of many exhibitions in China that need to be replaced, so that the exhibition can be applied persistently, the content and form of the exhibition stand can be replaced at will, and finally the intelligent level, interaction reliability and user experience of the exhibition stand are improved, which is more suitable for complex and diverse user and environmental needs in science and technology museums, museums and other scenes. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a flowchart of an AI interactive exhibition teaching method provided by an embodiment of the present application.
[0018] Figure 2 is a schematic diagram of an AI interactive exhibition teaching system provided by an embodiment of the present application. DETAILED DESCRIPTION
[0019] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be thoroughly understood and fully conveyed to those skilled in the art.
[0020] As Figure 1 shown, the embodiments of the present application propose an AI interactive exhibition learning booth intelligent control method, which comprises the following steps: Step 1, according to the weight sensor to perceive the user to come and start the process, through the camera to identify the user group attribute, load the adaptive exhibition learning content mode; Step 2, after loading the content mode, collect the environmental visual and audio information to obtain the initial environmental interference evaluation parameter; the camera takes three fixed physical structures on the booth as monitoring points, and the three monitoring points are the upper left corner of the interactive main screen, the upper right corner of the control panel and the central marker point of the front edge of the booth table surface, so as to obtain the preset marker area; Step 3, divide the preset marker area into multiple regular sub-areas; analyze the imaging feature change of each sub-area in the visual picture to obtain the environmental adjustment coefficient; calibrate the initial environmental interference evaluation parameter through the environmental adjustment coefficient to obtain the calibrated environmental interference evaluation parameter; Step 4, according to the calibrated environmental interference evaluation parameter, dynamically adjust the confidence threshold of gesture trajectory and semantic understanding to obtain the adjusted gesture trajectory confidence threshold and semantic understanding confidence threshold; Step 5, based on the adjusted gesture trajectory confidence threshold and semantic understanding confidence threshold, identify the user gesture or voice instruction; the voice instruction confidence exceeding the threshold is effective; the gesture instruction needs to meet the confidence threshold and pass the verification of the intersection judgment of the gesture trajectory and the preset line segment to confirm the effective instruction, so as to control the multimedia device to display the popular science content; Step 6, after the popular science content display is finished, analyze the user interaction behavior data to obtain the attention index; according to the attention index, optimize the loading priority of the exhibition learning content next time.
[0021] In the embodiment of the present application, through the cooperation of the weight sensor and the camera, automatic triggering of user arrival is realized, avoiding the problem that the traditional sensing method is easy to be blocked and mis-triggered, and the user group attribute can be accurately identified and the adaptive learning content mode can be loaded, solving the pain point that the existing exhibition content and user cognitive demand are not matched, leading to low interaction willingness; by collecting environmental visual and audio information to obtain initial interference parameters, and taking the fixed physical structure of the exhibition as the monitoring point to construct the preset marking area, combining the imaging feature change of the sub-area to generate the environmental adjustment coefficient to calibrate the interference parameters, the accuracy of environmental interference judgment is improved; based on the calibrated interference parameters, the confidence threshold of gesture and semantic understanding is dynamically adjusted, which can flexibly respond to different environmental interference intensities, reducing the misjudgment in high interference environment and the omission in low interference environment; the design that the voice instruction is effective when the threshold is exceeded, and the gesture instruction needs to meet the confidence threshold and pass the intersection verification of the preset line segment, further ensures the accuracy of instruction recognition, and ensures that the multimedia device can reliably control the display of popular science content; by analyzing the user interaction behavior data to obtain the attention index and optimizing the content loading priority next time, a closed loop from perception to interaction to optimization is formed, and finally the interaction reliability, content adaptability and user experience of the learning exhibition stand are comprehensively improved, which is more suitable for the actual application requirements of exhibition hall and other scenes.
[0022] In a preferred embodiment of the present application, step 1 can include: Step 1.1, receiving a trigger signal sent by a weight sensor arranged on the exhibition stand surface, the trigger signal being used to indicate that a user has arrived at the exhibition stand interaction area, specifically including: installing a weight sensor under the surface of the physical and mechanical science popular science exhibition stand in the science and technology museum, the weight sensor continuously monitors the pressure change borne by the exhibition stand surface, when 3 to 5 children walk to the exhibition stand interaction area and stand in front of the surface, the pressure value detected by the weight sensor reaches a pre-set trigger threshold value that can indicate that a user has arrived, at this time the weight sensor will send a trigger signal to the control unit of the exhibition stand, which clearly indicates that a user has arrived at the interaction area of the exhibition stand.
[0023] Step 1.2, after receiving the trigger signal, sending a start instruction to the camera of the exhibition stand, and receiving user image data collected by the camera, specifically including: after the control unit of the exhibition stand receives the trigger signal sent by the weight sensor, it will not be mis-started due to the blocking of surrounding tourists like the infrared sensing in the prior art, but will accurately send a start instruction to the high-definition camera installed above the exhibition stand, which can clearly capture the images of the users in the interaction area, the camera starts working immediately after receiving the start instruction, and captures the images of 3 to 5 children standing in the interaction area in real time, and then continuously transmits the captured user image data back to the control unit of the exhibition stand.
[0024] Step 1.3, group attribute recognition analysis is performed on the user image data, and the group attribute at least includes an age level and a group size to obtain the recognized group attribute, specifically including: after the control unit of the exhibition stand receives the user image data transmitted by the camera, the built-in group attribute recognition function is started, the face features, height proportions and other information of each child in the image are analyzed first, it is judged that the age of each child is in the corresponding age range of children, so that the age level of the group is determined as children, and then the number of users in the image is counted to determine that the size of the group is 3 to 5 people, and finally the recognized group attribute including the age level of children and the group size of 3 to 5 people is obtained.
[0025] Step 1.4, according to the recognized group attribute, the exhibition learning content mode adapted thereto is matched and loaded from the pre-set standardized exhibition content library; wherein, based on the AI control technology, the built-in question and answer engine is called and the display content is optimized in combination with the visual AI analysis result to drive the multimedia equipment and the standardized exhibition stand to dynamically switch and control, realize the intelligent interaction without the participation of traditional teachers and adapt to the change of exhibition theme, the AI interaction system can be divided into voice interaction system and equipment, and the display content description and AI question and answer in the drawer. Specifically including: first, after the control unit of the exhibition stand obtains the recognized group attribute, it will immediately access the standardized exhibition content library pre-stored in the local database of the exhibition stand. The fixed matching relationship between different age levels, different group sizes and corresponding exhibition learning content modes has been established in the exhibition learning content library in advance, wherein the exhibition learning content mode for children group mainly includes animation demonstration content of lever principle and content explained in simple language, rather than formula derivation content designed for adult group in the prior art; the control unit will accurately match in the exhibition learning content library according to the recognized group attribute of age level of children and group size of 3 to 5 people, select the exhibition learning content mode completely adapted to the group attribute from the library, after the matching is completed, the control unit will transmit the adapted exhibition learning content mode to the interactive main screen of the exhibition stand and complete the loading operation, so that the loaded exhibition learning content can be normally displayed on the interactive main screen for 3 to 5 children standing in the interactive area to watch.
[0026] Based on the AI control technology, the control unit of the exhibition stand can call the built-in question and answer engine, and optimize the display content loaded on the interactive main screen in combination with the visual AI analysis result. The visual AI analysis result can feed back the concentration of the child when watching the content in real time, and the question and answer engine can prepare the adapted answer content in advance according to the possible doubts of the child, so as to ensure that the display content meets the cognitive level of the child and responds to the potential interactive demand in time through the combination of the two. The optimized display content can drive the multimedia equipment and the standardized exhibition to switch and control dynamically. For example, when the animation of the principle of the lever reaches the key node, the multimedia equipment can automatically adjust the picture brightness to highlight the key points, and the standardized exhibition can cooperate with the animation demonstration to show the corresponding physical structure. The whole process does not need the traditional teacher to participate in the operation, and can automatically adjust the display content and the state of the exhibition according to the change of the exhibition theme, so as to realize the intelligent interaction adapting to the change of the exhibition theme.
[0027] In the embodiment of the present application, because the weight sensor arranged on the exhibition stand is used to receive the trigger signal to indicate that the user arrives, the camera is started to collect the image data of the user and identify the age level and the group attribute of the group size after receiving the trigger signal, the built-in question and answer engine is called based on the AI control technology, the display content is optimized in combination with the visual AI analysis result, and the multimedia equipment and the standardized exhibition are driven to switch and control dynamically, so that the technical problems of the traditional sensing mode being easily blocked or mis-triggered, the exhibition content not matching the user group attribute, the need for traditional teachers to participate and the difficulty in adapting to the change of the exhibition theme are overcome, and the technical effects of accurately starting the user interactive process, loading the exhibition content mode adapted to the user group attribute, realizing the intelligent interaction without the participation of traditional teachers and adapting to the change of the exhibition theme are achieved.
[0028] In a preferred embodiment of the present application, the above step 2 can include: Step 2.1, after completing the loading of the exhibition content mode, the camera and audio acquisition device are instructed to collect the visual picture and audio information of the current environment; based on the strength and characteristics of the collected information, the initial environmental interference evaluation parameters are obtained, which specifically include: after the completion of the adaptation of the 3 to 5 children group's lever principle animation demonstration and simple language explanation content mode at the science museum physics and mechanics popular science exhibition stand, the control unit of the stand sends data collection instructions to the camera installed above the stand and the audio acquisition device distributed around the stand. The camera immediately starts capturing the visual picture of the current stand, focusing on recording the brightness anomaly, screen reflection and surrounding tourists walking images generated by the afternoon sunlight directly hitting the interactive main screen; the audio acquisition device synchronously collects the noise generated by the crowd in the exhibition hall, the sound of other exhibition equipment running and the conversation between children. After receiving the visual picture and audio information, the brightness intensity, reflection area characteristics of the visual picture and the noise intensity, sound frequency characteristics of the audio information are fused and calculated, such as combining the degree of picture brightness exceeding the normal range caused by strong light and the decibel value of the noise, to finally obtain the initial environmental interference evaluation parameters that can preliminarily reflect the current visual and audio interference degree.
[0029] Step 2.2, based on the initial environmental interference evaluation parameters, the environmental interference is analyzed, and the pre-stored stand physical structure coordinate data is called, which accurately corresponds to the pre-set positions of the upper left corner of the interactive main screen, the upper right corner of the control panel and the central marker point of the front edge of the stand table in the camera picture, specifically including: based on the obtained initial environmental interference evaluation parameters, the interference type and interference intensity of the current environment are analyzed, and it is determined that there is visual interference caused by afternoon sunlight and audio interference caused by crowd noise in the exhibition hall. In order to accurately lock the interference of the core interaction area of the stand and exclude the interference of the surrounding irrelevant area, the control unit calls the stand physical structure coordinate data pre-stored in the local database. The coordinate data is obtained by camera calibration before the stand is factory-finished, which accurately corresponds to the specific position of the upper left corner of the interactive main screen in the camera picture, the specific position of the upper right corner of the control panel in the camera picture and the specific position of the central marker point of the front edge of the stand table in the camera picture, ensuring that each coordinate accurately points to the key physical structure of the children's interactive operation in the stand, and is irrelevant to the tourist activity area outside the stand.
[0030] Step 2.3, based on the preset position coordinates of the three monitoring points obtained, a triangular region covering the three monitoring points is constructed by taking the coordinates as vertices through geometric operation, and finally the triangular region is defined as the preset marked region, specifically including: after obtaining the preset position coordinates of the three monitoring points of the upper left corner of the interactive main screen, the upper right corner of the control panel and the central mark point at the front edge of the exhibition table in the camera picture, taking the three coordinates as the three vertices of the triangle respectively, the three vertices are sequentially connected to form a closed figure through geometric operation, and a triangular region capable of completely covering the three monitoring points is constructed, the triangular region covers the exhibition table area where the children perform gesture operation, the interactive main screen area where the children watch popular science content and the control panel area where the children perform auxiliary operation, which belongs to the core interaction area of the exhibition table, and finally the control unit defines the triangular region as the preset marked region.
[0031] In the embodiment of the application, after the exhibition content mode loading is completed, the camera and the audio acquisition device are instructed to collect the visual picture and the audio information of the current environment, and the initial environment interference evaluation parameter is calculated based on the strength and feature fusion of the collected information; then, based on the initial environment interference evaluation parameter, the environment interference is analyzed, the pre-stored exhibition table physical structure coordinate data accurately corresponding to the preset positions of the upper left corner of the interactive main screen, the upper right corner of the control panel and the central mark point at the front edge of the exhibition table in the camera picture are called; finally, the preset position coordinates of the three monitoring points are taken as the vertices, a triangular region covering the three monitoring points is constructed through geometric operation, and the triangular region is defined as the preset marked region, which overcomes the technical problems that the existing interactive exhibition table lacks an interference calibration mechanism based on real-time environmental features, and does not focus on the core interaction area of the exhibition table, resulting in inaccurate environmental interference judgment, and thus realizes accurate capture of the initial interference situation of the current environment, and through the preset marked region focusing on the key interaction area of the exhibition table and excluding the invalid interference of the surrounding irrelevant area, the pertinence and accuracy of the environmental interference evaluation are improved.
[0032] In a preferred embodiment of the application, the above-mentioned step 3 can include: Step 3.1, after obtaining the preset marking area, the triangular area is divided into a plurality of rectangular sub-areas with equal area according to the predefined grid rule, specifically including: after obtaining the triangular preset marking area covering the upper left corner of the interactive main screen, the upper right corner of the control panel and the central marking point of the front edge of the exhibition stand, the exhibition stand control unit divides the triangular area according to the predefined grid rule, which is set to divide the triangular area into three equal parts horizontally and two equal parts vertically. By this division method, the original triangular area is divided into a plurality of rectangular sub-areas with equal area. The rectangular sub-areas include the interactive main screen part for children to watch the animation of the principle of levers, cover the control panel area that children may touch and operate, and involve the exhibition stand surface area for children to stand and interact.
[0033] Step 3.2, based on the divided sub-areas, the imaging features of each sub-area in the continuous time sequence of visual pictures are extracted in turn, including average brightness, color distribution and pixel motion vector, and the change amount of each sub-area imaging feature relative to the initial reference frame is calculated to obtain the statistical result of the imaging feature change amount, specifically including: taking each divided rectangular sub-area as an independent analysis object, the imaging features of each sub-area in the continuous time sequence of visual pictures are extracted in turn. The continuous time sequence is set to continuously collect fifteen frames of visual pictures, and the first frame is taken as the initial reference frame. For each sub-area, the average brightness feature is first extracted, and the sub-area above the interactive main screen directly irradiated by the afternoon sun is monitored. The average brightness of the area is higher than that of other sub-areas not directly irradiated. Then the color distribution feature is extracted, and the deviation of the color in the sub-area caused by the direct irradiation of the sun is observed, such as the yellowish color of the originally green animation elements under direct irradiation. Then the pixel motion vector feature is extracted, and the pixel movement trajectory in the sub-area near the exhibition stand caused by the passing visitors in the exhibition hall is captured. After extracting the imaging features of each sub-area in each frame, the control unit calculates the difference between the imaging features of each frame and the corresponding features of the initial reference frame to obtain the change amount of each sub-area in average brightness, color distribution and pixel motion vector. Finally, the change amount is summarized and counted to form the statistical result of the imaging feature change amount of each sub-area, such as the average brightness change amount of a sub-area directly irradiated by the sun reaching twenty-five percent, and the pixel motion vector change amount of a sub-area near the exhibition hall corridor being significantly higher than that of other sub-areas.
[0034] Step 3.3, based on the statistical results of the imaging feature variation of all sub-regions, a comprehensive environmental adjustment coefficient is calculated by weighted fusion; the environmental adjustment coefficient is multiplied by the obtained initial environmental disturbance evaluation parameter to obtain the calibrated environmental disturbance evaluation parameter, which specifically includes: based on the statistical results of the imaging feature variation of all rectangular sub-regions, a comprehensive environmental adjustment coefficient is calculated by weighted fusion, and different imaging feature variation weights are set according to the influence degree of the environmental disturbance in the exhibition hall on the interaction, wherein the weight of the average brightness variation is the highest, because the direct sunlight in the afternoon is the most prominent disturbance to visual identification, the color distribution variation weight is the second, and the pixel motion vector variation weight is relatively low. The control unit performs fusion operation on the imaging feature variation of all sub-regions according to the set weight, and finally obtains the comprehensive environmental adjustment coefficient, and then the control unit multiplies the environmental adjustment coefficient with the initial environmental disturbance evaluation parameter calculated by fusing the visual picture and the audio information. Since the initial environmental disturbance evaluation parameter does not fully consider the local disturbance difference of each sub-region, such as not accurately reflecting the strong disturbance of the sunlight direct sunlight sub-region, after the multiplication operation, the calibrated environmental disturbance evaluation parameter can more truly reflect the actual disturbance of the core interactive area of the exhibition stand.
[0035] In the embodiment of the present application, after obtaining the preset marked area, the triangular area is divided into a plurality of rectangular sub-regions with equal area according to the pre-defined grid rule; based on the divided sub-regions, the average brightness, color distribution and pixel motion vector of each sub-region in the continuous time sequence visual picture are extracted in turn, the variation of the imaging feature of each sub-region relative to the initial reference frame is calculated and the statistical results are obtained; and based on the statistical results of the imaging feature variation of all sub-regions, a comprehensive environmental adjustment coefficient is calculated by weighted fusion, and the environmental adjustment coefficient is multiplied by the initial environmental disturbance evaluation parameter to obtain the calibrated environmental disturbance evaluation parameter. The technical means overcomes the technical problem that the existing interactive exhibition learning exhibition stand only stays at the overall general level in evaluating the environmental disturbance, does not perform subdivision analysis on the core interactive area, cannot accurately capture the local disturbance in the area, and the initial environmental disturbance evaluation parameter deviates greatly from the actual disturbance, and further realizes accurate capture of the disturbance characteristics of each local sub-region in the preset marked area, so that the environmental adjustment coefficient can fully reflect the real disturbance details in the area, and finally the calibrated environmental disturbance evaluation parameter is more consistent with the actual disturbance state of the core interactive area of the exhibition stand.
[0036] In a preferred embodiment of the present application, the above step 4 can include: Step 4.1, based on the calibrated environmental interference evaluation parameter, query the pre-stored mapping relationship table; the mapping relationship table defines the corresponding relationship between the environmental interference evaluation parameter value and the confidence threshold adjustment amount, so as to respectively obtain the basic adjustment amount of gesture trajectory confidence and the basic adjustment amount of semantic understanding confidence, specifically including: based on the calibrated environmental interference evaluation parameter, the audio interference of the crowd noise in the exhibition hall and the visual interference of the direct sunlight in the afternoon are integrated, the value is at a high level, the mapping relationship table pre-stored in the local database is called, and the mapping relationship table defines the gesture trajectory confidence adjustment amount and the semantic understanding confidence adjustment amount corresponding to different environmental interference evaluation parameter values. Since the current calibrated parameter shows that the environmental interference is strong, the control unit queries the corresponding gesture trajectory confidence basic adjustment amount from the mapping relationship table, which is a positive value for improving the gesture recognition threshold to reduce misjudgment; at the same time, the semantic understanding confidence basic adjustment amount is queried, which is a proper positive value, considering the need to improve the threshold due to the interference of the crowd noise, and avoiding too high threshold due to the recognition of children's voices.
[0037] Step 4.2, superimpose the basic adjustment amount of gesture trajectory confidence and a preset gesture trajectory confidence basic threshold to obtain an adjusted gesture trajectory confidence threshold, specifically including: after obtaining the basic adjustment amount of gesture trajectory confidence, superimpose the gesture trajectory confidence and the preset gesture trajectory confidence basic threshold, the preset basic threshold is a reference value that can accurately recognize the effective gestures of children in the absence of obvious environmental interference, and the current environment is prone to gesture recognition misjudgment due to direct sunlight and crowd movement, so the basic adjustment amount is positive. Through superimposition operation, the adjusted gesture trajectory confidence threshold is obtained.
[0038] Step 4.3, superimpose the basic adjustment amount of semantic understanding confidence and a preset semantic understanding confidence basic threshold to obtain an adjusted semantic understanding confidence threshold, specifically including: after obtaining the basic adjustment amount of semantic understanding confidence, superimpose the semantic understanding confidence and the preset semantic understanding confidence basic threshold, the preset basic threshold is a reference value that can accurately recognize the children's voice commands in a quiet environment, and the current environment is disturbed by the crowd noise, but the basic adjustment amount is set to a proper positive value, which is higher than the basic threshold to filter part of the noise, and does not completely shield the children's voice like the fixed high threshold in the prior art. The adjusted semantic understanding confidence threshold obtained by superimposition operation can accurately capture the effective voice commands of children explaining the principle of levers in a crowd noisy environment, solving the problem that the fixed high threshold in the prior art cannot respond to the children's voice commands.
[0039] In the embodiment of the present application, the preset mapping relationship table is queried based on the calibrated environmental interference evaluation parameter, and the basic adjustment amount of gesture track confidence and the basic adjustment amount of semantic understanding confidence are obtained respectively, and then the basic adjustment amount of gesture track confidence is superimposed with the preset gesture track confidence basic threshold, and the basic adjustment amount of semantic understanding confidence is superimposed with the preset semantic understanding confidence basic threshold, to obtain the adjusted gesture track confidence threshold and the adjusted semantic understanding confidence threshold respectively. The technical means solves the technical problem that the existing interactive exhibition stands adopt the factory-preset fixed gesture and voice recognition threshold, which cannot be dynamically adjusted according to the real-time environmental interference such as the crowd noise in the exhibition hall and the direct sunlight in the afternoon, resulting in that the preset low gesture threshold is easy to misjudge the children's random waving as a content switching instruction in a high interference environment, and the preset high voice threshold cannot always respond to the child's voice instruction of explaining the principle of the lever, and it is difficult to balance the interaction accuracy and sensitivity in different environments, and then the precise adaptation of the gesture track and semantic understanding confidence threshold to the real-time environmental interference state is realized, so that the misjudgment is reduced and the omission is avoided in a high interference environment, and the interaction sensitivity is ensured in a low interference environment.
[0040] In a preferred embodiment of the present application, the above step 5 can include: Step 5.1, based on the obtained adjusted gesture track confidence threshold and semantic understanding confidence threshold, the gesture of the real-time video stream collected by the camera and the voice of the real-time audio stream collected by the audio collection device are recognized and analyzed to obtain real-time gesture recognition results and voice recognition analysis results, specifically including: after determining the adjusted gesture track confidence threshold and the semantic understanding confidence threshold, the stand control unit starts the camera to capture the actions of 3 to 5 children in the interactive area in real time, generates a real-time video stream, and simultaneously starts the audio collection device to collect the voices of the children and the environmental sounds in the exhibition hall, and generates a real-time audio stream. The control unit recognizes and analyzes the children's gestures in the real-time video stream based on the adjusted gesture track confidence threshold, and distinguishes between the children's random waving and the interactive gestures intentionally made, such as pointing to the screen and horizontal sliding, to obtain real-time gesture recognition results containing gesture track confidence. At the same time, based on the adjusted semantic understanding confidence threshold, the voice in the real-time audio stream is recognized and analyzed, the noise generated by the crowd noise is filtered, and the voice of the children such as explaining the principle of the lever and looking again is captured, to obtain real-time voice recognition analysis results containing voice instruction confidence.
[0041] Step 5.2, when the confidence of a certain voice instruction in the voice recognition analysis result exceeds the adjusted semantic understanding confidence threshold, it is determined that the voice instruction is a valid instruction, specifically including: checking each voice instruction confidence in the real-time voice recognition analysis result one by one, when it is detected that the confidence of a certain child's explanation of the principle of the lever voice instruction exceeds the adjusted semantic understanding confidence threshold, it is immediately determined that the voice instruction is a valid instruction. In this process, the adjusted threshold not only filters the interference of the crowd noise in the exhibition hall, but also does not shield the child's voice instruction due to the high threshold, solving the problem that the fixed high threshold in the prior art cannot respond to the child's voice instruction.
[0042] Step 5.3, judging whether the confidence of the gesture trajectory in the gesture recognition result exceeds the adjusted gesture trajectory confidence threshold; if it exceeds, further performing geometric intersection judgment on the motion trajectory of the gesture and a virtual line segment preset in the interaction space to obtain a geometric intersection judgment result, specifically including: judging whether the confidence of each gesture trajectory in the real-time gesture recognition result exceeds the adjusted gesture trajectory confidence threshold, when a child makes a horizontal sliding gesture, and the confidence of the horizontal sliding gesture trajectory exceeds the threshold, the control unit further calls the virtual line segment data preset in the interaction space. The virtual line segment is a horizontal line segment parallel to the bottom edge of the interactive main screen and located in the child's gesture operation range. The control unit judges whether the motion trajectory of the child's horizontal sliding gesture intersects with the virtual line segment through geometric operation, and further obtains the geometric intersection judgment result.
[0043] Step 5.4, according to the geometric intersection judgment result, when and only when the motion trajectory of a certain gesture instruction intersects with the preset virtual line segment, finally determining that the gesture instruction is a valid instruction, specifically including: further determining according to the geometric intersection judgment result, if the motion trajectory of the child's gesture intersects with the preset virtual line segment, it means that the gesture is a valid operation made for the exhibition stand interaction area, and finally determining that the gesture instruction is a valid instruction; if the motion trajectory of the child's gesture does not intersect with the preset virtual line segment, such as the child waving his hand randomly on the side of his body, it is determined that the gesture instruction is invalid. Through the judgment, the child's non-interactive intention gesture is excluded, and the problem that the random waving hand is misjudged as a valid instruction due to the low threshold in the prior art is avoided.
[0044] Step 5.5, after determining any valid instruction, generate a control signal corresponding to the valid instruction, and send it to the multimedia devices of the exhibition stand; wherein the control signal is used to drive the multimedia devices to display the popular science content corresponding to the instruction, and to dynamically adjust the extension and lifting of the display screen and the opening and closing of the interactive multifunctional drawer according to the needs of the current exhibition theme, to realize optimized space utilization and audience interaction effect by combining the control of the exhibition equipment, specifically including: when the control unit of the exhibition stand determines a valid instruction, it will first analyze the specific content of the valid instruction, such as a voice instruction to explain the principle of lever issued by a child, or a horizontal sliding gesture instruction; then, the control unit will look up the control signal matching the valid instruction in the pre-established instruction and control signal corresponding relationship library. This corresponding relationship library stores various instructions and their corresponding specific control parameters, such as the signal parameters of driving the screen to play the video of explaining the principle of lever for the instruction to explain the principle of lever, and the signal parameters of switching to the next demonstration content for the horizontal sliding gesture instruction. The control unit generates a control signal containing specific execution parameters according to the matching relationship found; then, the control unit sends the generated control signal to the multimedia devices of the exhibition stand through the internal communication line of the exhibition stand, including the interactive main screen, the extension and lifting driving device of the display screen, and the opening and closing control device of the interactive multifunctional drawer, etc. After receiving the control signal, the multimedia devices will perform the corresponding operation according to the parameters in the signal. If the valid instruction is to explain the principle of lever, the interactive main screen will immediately call up and play the corresponding popular science video of the principle of lever, and the audio device will simultaneously play clear explanation sound; if the valid instruction is horizontal sliding, the main screen will switch to the next relevant demonstration content.
[0045] Meanwhile, according to the specific needs of the current exhibition theme, the control signal will drive the extension and lifting device of the display screen to start under the guidance of the exhibition structure layout and mode of the natural center "natural cognitive reconstruction" concept and method, for example, when the exhibition theme emphasizes children's hands-on operation, the display screen will automatically shrink inward by a certain distance and lower the height, so that children can more easily see the screen content and touch the operation area; if the exhibition theme needs to show more content, the display screen will extend outward and rise, expanding the display range to present more information. The opening and closing of the interactive multifunctional drawer will also be adjusted according to the control signal. For example, when the actual assembly part of the lever is explained, the corresponding drawer will automatically open, exposing the inside lever assembly and operation tools, making it convenient for children to practice; when the demonstration is over or switched to other content, the drawer will automatically close, keeping the exhibition table surface clean and preventing components from being lost. Through such combination control of multimedia equipment and exhibition equipment, not only can the popular science content corresponding to the effective instruction be accurately and timely displayed, but also the space state of the exhibition equipment can be dynamically adjusted according to the exhibition theme, optimizing the space utilization efficiency of the exhibition table, making the interactive process more convenient and comfortable for children, and improving the overall audience interaction effect.
[0046] In the embodiment of the present application, based on the adjusted gesture track confidence threshold and semantic understanding confidence threshold, the real-time video stream gesture collected by the camera and the real-time audio stream voice collected by the audio collection device are recognized and analyzed, the voice instruction confidence is determined as an effective instruction when it exceeds the threshold, and the gesture instruction is determined as an effective instruction when it exceeds the threshold and is geometrically intersected with the preset virtual line segment in the interaction space. At the same time, after determining the effective instruction, the corresponding control signal is generated and sent to the multimedia device to drive it to display the corresponding popular science content, and the display screen extension and lifting and the interactive multifunctional drawer opening and closing are dynamically adjusted according to the exhibition theme. Therefore, the technical problems of voice miss or gesture misjudgment in high interference environment caused by fixed recognition threshold in traditional exhibition table, lack of secondary verification of gesture recognition leading to insufficient accuracy, disconnection between exhibition equipment control and content display, and low space utilization efficiency are overcome, and the technical effects of accurately recognizing user's effective instruction in high interference environment, reliably controlling multimedia equipment to display matching popular science content, optimizing space utilization and audience interaction effect through exhibition equipment combination control, and improving exhibition interaction accuracy and user experience are achieved.
[0047] In a preferred embodiment of the present application, step 6 can include: Step 6.1, after the end of the science content display, collect the user interaction behavior data in this interaction process, specifically including: after completing the science content display for 3 to 5 children groups at the physical mechanics science popularization exhibition stand in the science museum, i.e. the children have watched the lever principle animation explanation, lever application example demonstration and other content and left the interaction area, and the weight sensor detects that the pressure value falls back to the no user state, the stand control unit automatically collects the user interaction behavior data in this interaction process. The collected data includes the children's viewing time of each segment of science content, such as 8 minutes for watching the lever principle animation, 5 minutes for watching the lever application example, and only 1 minute for skipping the formula derivation content; also includes the number of times the children trigger various instructions, such as issuing the look again voice instruction twice for the lever principle animation, and triggering the enlarge animation detail instruction 3 times by gesture; at the same time, it includes whether the children trigger the content detail viewing operation, such as 2 children clicking to view the details page of the application of lever principle in life, and no detail operation of formula derivation content is triggered.
[0048] Step 6.2, based on the collected user interaction behavior data, the attention index for quantifying the user's interest degree in different science content is obtained by comprehensive calculation through the weight calculation formula, specifically including: after obtaining the collected user interaction behavior data, the attention index is obtained by comprehensive calculation according to the weight calculation formula. In the weight calculation formula, the weight of viewing time is set to the highest, because the longer the viewing time, the more it reflects the children's interest in the content; the instruction trigger times weight is second, reflecting the children's initiative to interact; the content detail viewing operation weight is relatively low, as a supplementary reference for interest degree. After comparing the viewing time of each segment of science content with the average viewing time of all content, comparing the instruction trigger times with the total instruction times, and assigning values after detail viewing operation, the set weights are combined for superposition operation. For example, the viewing time of the lever principle animation is much longer than the average level, the instruction trigger times are the most and there is detail viewing operation, the calculated attention index is 1.8; the attention index of the lever application example is 1.2; the formula derivation content has short viewing time, no instruction trigger and no detail viewing, and the attention index is only 0.3.
[0049] Step 6.3, according to the calculated attention index, the loading priority weight of the corresponding science popularization content in the content library is dynamically updated to obtain the updated loading priority weight, specifically including: according to the calculated attention index of each science popularization content, the loading priority weight of the corresponding science popularization content in the local content library is dynamically updated, the original loading priority weight of each science popularization content in the content library is 1.0, and the content is loaded in a fixed order, which may cause the formula derivation content that children are not interested in to be displayed first. Now the control unit adjusts the weight according to the attention index, updates the loading priority weight of the lever principle animation from 1.0 to 1.8, updates the weight of the lever application instance to 1.2, and updates the weight of the formula derivation content to 0.3, so that the content with higher attention index has higher loading priority weight.
[0050] Step 6.4, based on the updated loading priority weight, combined with the display device adopting a standardized interface, the optimized content loading order is persistently stored in the database; wherein the standardized display device supports flexible replacement of display content and physical form, so that the display stand adapts to different exhibition themes according to the updated priority, realizes the persistent application of the exhibition, specifically including: the display stand control unit first obtains the updated loading priority weight of each science popularization content, sorts the science popularization content in the content library according to the weight from high to low, determines the optimized content loading order, for example, the lever principle animation is ranked first, the lever application instance is ranked second, and the formula derivation content is ranked last, forming the exclusive content loading order for the children group after this interaction; then, the control unit combines the characteristics of the display device adopting the standardized interface, and sorts the corresponding relationship between each science popularization content and the physical form of the display device. The standardized display device includes an interactive main screen, an interactive multifunctional drawer, and a physical demonstration component, etc., and its interface follows a unified technical specification, supporting flexible replacement of display content and physical form, for example, the lever principle animation needs to be displayed through the interactive main screen, and the corresponding display device interface needs to ensure that the animation signal can be stably transmitted to the screen; the lever physical component needs to be presented through the interactive multifunctional drawer, and the interface needs to control the opening and closing time sequence and angle of the drawer to match the content display rhythm.
[0051] The control unit integrates the optimized content loading order and corresponding exhibit control parameters, such as screen display resolution and drawer opening and closing speed, into structured data, and the data clearly marks the loading priority of each item of content, the corresponding exhibit type and specific control instructions. Then, the control unit stores the structured data persistently in the local database of the exhibition stand through a data writing interface, and starts a data verification mechanism to check whether the stored data is consistent with the original order and control parameters, ensuring the accuracy of data storage and avoiding abnormal loading order due to data loss or errors in the next loading; when the subsequent exhibition theme changes, for example, from the principle of levers to the principle of pulleys, since the exhibits use standardized interfaces, there is no need to modify the hardware of the exhibits, only the popular science content related to the principle of pulleys needs to be updated in the content library, and the optimized loading priority weight calculation logic is used. The control unit reads the persistently stored loading order rules from the database when it starts next time, combines the content weight under the new theme, and automatically adapts the physical display form of the exhibits, such as preferentially loading the pulley principle animation to the interactive main screen and presenting the pulley physical components through the interactive multifunctional drawer, to realize the persistent application of the exhibition stand under different exhibition themes without frequent replacement of exhibit hardware, thereby reducing operating costs.
[0052] In the embodiment of the present application, the technical means of collecting user interaction behavior data after the popular science content display is completed, calculating the attention index of the quantified user interest degree in different popular science contents based on the data through a weight calculation formula, dynamically updating the loading priority weight of the corresponding popular science content in the content library according to the attention index, and persistently storing the optimized content loading order to the database in combination with the exhibits using standardized interfaces are adopted, so that the technical problems of fixed loading order of interactive exhibition and learning exhibition content, lack of personalized optimization mechanism based on actual user interaction behavior, inability to dynamically adjust content priority according to user interest, and difficulty of adapting different exhibition themes for exhibits are overcome, thereby achieving the technical effects of dynamically optimizing the loading priority of exhibition and learning content according to user interaction behavior, preferentially loading the content with high attention degree of the user when the same group attribute user is recognized next time, and adapting different exhibition themes for the standardized exhibits to realize persistent application of the exhibition, and improving the matching degree of content and user interest and the practicality of the exhibition and learning exhibition stand.
[0053] As shown in Figure 2 The embodiment of the present application also provides an AI interactive exhibition and learning exhibition stand intelligent control system, which comprises: The acquisition module is used for starting the process by sensing the arrival of a user through a weight sensor, identifying the group attribute of the user through a camera, and loading the adapted exhibition and learning content mode. An adjusting module is configured to collect environmental visual and audio information after the content mode is loaded to obtain initial environmental interference evaluation parameters; the camera takes three fixed physical structures on the exhibition stand as monitoring points, and the three monitoring points are the upper left corner of the interactive main screen, the upper right corner of the control panel, and the central marker point of the front edge of the exhibition stand surface to obtain a preset marker area; the preset marker area is divided into multiple regular sub-areas; imaging feature changes of each sub-area in the visual picture are analyzed to obtain an environmental adjustment coefficient; the initial environmental interference evaluation parameters are calibrated by the environmental adjustment coefficient to obtain calibrated environmental interference evaluation parameters; A computing module is configured to dynamically adjust a gesture trajectory and a semantic understanding confidence threshold according to the calibrated environmental interference evaluation parameters to obtain an adjusted gesture trajectory confidence threshold and a semantic understanding confidence threshold; a user gesture or a voice instruction is identified based on the adjusted gesture trajectory confidence threshold and the semantic understanding confidence threshold; the voice instruction confidence exceeding the threshold is effective; the gesture instruction needs to simultaneously satisfy the confidence threshold and pass the verification of the intersection judgment of the gesture trajectory and the preset line segment to confirm the effective instruction, thereby controlling the multimedia device to display the science popularization content; A processing module is configured to analyze user interaction behavior data to obtain an attention index after the science popularization content display is finished; and the loading priority of the exhibition content is optimized according to the attention index when starting next time.
[0054] The above is the preferred embodiment of the present application. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should also be considered within the scope of protection of the present application.
Claims
1. An intelligent control method for AI interactive educational exhibition booths, characterized in that: The method includes: The system detects the arrival of users using a weight sensor and initiates the process, identifies user group attributes using a camera, and loads appropriate educational content modes. After the content mode is loaded, environmental visual and audio information is collected to obtain initial environmental interference assessment parameters; the camera uses three fixed physical structures on the exhibition stand as monitoring points, namely the upper left corner of the interactive main screen, the upper right corner of the control panel, and the central marker point on the front edge of the exhibition stand, to obtain the preset marking area. The preset marked area is divided into multiple regular sub-regions; the imaging feature changes of each sub-region in the visual image are analyzed to obtain the environmental adjustment coefficient; the initial environmental interference assessment parameters are calibrated using the environmental adjustment coefficient to obtain the calibrated environmental interference assessment parameters. Based on the calibrated environmental interference assessment parameters, the confidence thresholds for gesture trajectory and semantic understanding are dynamically adjusted to obtain the adjusted confidence thresholds for gesture trajectory and semantic understanding. Based on the adjusted confidence thresholds for gesture trajectory and semantic understanding, user gestures or voice commands are recognized; voice commands become effective when their confidence exceeds the threshold; gesture commands must simultaneously meet the confidence threshold and be verified by judging the intersection of the gesture trajectory with a preset line segment to confirm the validity of the command, thereby controlling the multimedia device to display popular science content. After the science popularization content presentation ends, user interaction data is analyzed to obtain attention metrics; based on these attention metrics, the loading priority of the educational content is optimized for the next launch.
2. The intelligent control method for the AI interactive educational exhibition booth according to claim 1, characterized in that, The system detects user arrival using a weight sensor and initiates the process, identifies user group attributes using a camera, and loads appropriate educational content modes, including: It receives a trigger signal from a weight sensor located on the booth surface, which indicates that a user has arrived at the booth's interactive area. Upon receiving the trigger signal, it sends a start command to the camera at the booth and receives user image data captured by the camera. Perform group attribute identification and analysis on user image data. The group attributes include at least age level and group size in order to obtain the identified group attributes. Based on the identified group attributes, the system matches and loads appropriate exhibition and learning content modes from a pre-set standardized exhibition equipment content library. Among these features, based on AI control technology, the system calls upon the built-in question-and-answer engine and combines visual AI analysis results to optimize the display content, thereby driving the multimedia equipment and standardized exhibition equipment to dynamically switch and control, achieving intelligent interaction that does not require traditional teacher involvement and adapts to changes in the exhibition theme.
3. The intelligent control method for the AI interactive educational exhibition booth according to claim 2, characterized in that, After the content mode is loaded, environmental visual and audio information is collected to obtain initial environmental interference assessment parameters. The camera uses three fixed physical structures on the exhibition stand as monitoring points: the upper left corner of the interactive main screen, the upper right corner of the control panel, and the central marker point on the front edge of the exhibition stand, to obtain the preset marking area, including: After loading the learning content mode, commands are sent to the camera and audio acquisition device to collect visual images and audio information of the current environment; based on the intensity and characteristics of the collected information, fusion calculations are performed to obtain initial environmental interference assessment parameters; Based on the initial environmental interference assessment parameters, environmental interference analysis is performed, and the pre-stored physical structure coordinate data of the booth is called. The coordinate data precisely corresponds to the preset positions of the upper left corner of the interactive main screen, the upper right corner of the control panel, and the center marker point of the front edge of the booth in the camera image. Based on the preset location coordinates of the three monitoring points, a triangular region covering the three monitoring points is constructed through geometric operations using the coordinates as vertices. Finally, the triangular region is defined as the region marked by the preset coordinates.
4. The intelligent control method for the AI interactive educational exhibition booth according to claim 3, characterized in that, The preset marked area is divided into multiple regular sub-regions; the imaging feature changes of each sub-region in the visual image are analyzed to obtain the environment adjustment coefficient; The initial environmental disturbance assessment parameters are calibrated using an environmental adjustment factor to obtain the calibrated environmental disturbance assessment parameters, including: After obtaining the preset marked area, the triangular area is divided into multiple rectangular sub-regions of equal area according to the predefined grid rules; Based on the divided sub-regions, the imaging features of each sub-region in the visual image of the continuous time series are extracted in turn. The imaging features include average brightness, color distribution and pixel motion vectors. The change of the imaging features of each sub-region relative to the initial reference frame is calculated to obtain the statistical results of the change of imaging features. Based on the statistical results of the changes in imaging features of all sub-regions, a comprehensive environmental adjustment coefficient is calculated by weighted fusion. The environmental adjustment coefficient is then multiplied by the obtained initial environmental interference assessment parameters to obtain the calibrated environmental interference assessment parameters.
5. The intelligent control method for the AI interactive educational exhibition booth according to claim 4, characterized in that, Based on the calibrated environmental interference assessment parameters, the confidence thresholds for gesture trajectory and semantic understanding are dynamically adjusted to obtain the adjusted confidence thresholds for gesture trajectory and semantic understanding, including: Based on the calibrated environmental interference assessment parameters, a pre-set mapping table is queried; the mapping table defines the correspondence between the environmental interference assessment parameter values and the confidence threshold adjustment amount, thereby obtaining the basic adjustment amount of the gesture trajectory confidence and the basic adjustment amount of the semantic understanding confidence respectively; The adjusted confidence threshold of the gesture trajectory is obtained by superimposing the basic adjustment amount of the confidence of the gesture trajectory with a preset basic confidence threshold of the gesture trajectory. The adjusted semantic understanding confidence threshold is obtained by superimposing the basic adjustment amount of the semantic understanding confidence with a preset basic threshold of semantic understanding confidence.
6. The intelligent control method for the AI interactive educational exhibition booth according to claim 5, characterized in that, Based on the adjusted gesture trajectory confidence threshold and semantic understanding confidence threshold, user gestures or voice commands are recognized; voice commands become effective when their confidence exceeds the threshold; gesture commands must simultaneously meet the confidence threshold and pass verification by judging the intersection of the gesture trajectory with a preset line segment to confirm a valid command, thereby controlling the multimedia device to display popular science content, including: Based on the adjusted gesture trajectory confidence threshold and semantic understanding confidence threshold, the gestures in the real-time video stream captured by the camera and the speech in the real-time audio stream captured by the audio acquisition device are recognized and analyzed to obtain real-time gesture recognition results and speech recognition analysis results. When the confidence level of a certain voice command in the speech recognition analysis results exceeds the adjusted semantic understanding confidence threshold, the voice command is determined to be a valid command. Determine whether the confidence level of the gesture trajectory in the gesture recognition result exceeds the adjusted gesture trajectory confidence level threshold; if it does, further perform a geometric intersection judgment between the gesture trajectory and a virtual line segment preset in the interaction space to obtain the result of the geometric intersection judgment. Based on the result of geometric intersection judgment, the gesture command is finally determined to be a valid command if and only if the motion trajectory of a certain gesture command intersects with the preset virtual line segment; After determining any valid instruction, a control signal corresponding to the valid instruction is generated and sent to the multimedia equipment on the exhibition stand. The control signal is used to drive the multimedia equipment to display the popular science content corresponding to the instruction, and to dynamically adjust the extension and retraction of the display screen and the opening and closing of the interactive multi-functional drawer according to the needs of the current exhibition theme. By combining and controlling the exhibition equipment, optimized space utilization and audience interaction effects can be achieved.
7. The intelligent control method for the AI interactive educational exhibition booth according to claim 6, characterized in that, After the science popularization content presentation, user interaction data was analyzed to obtain attention metrics; Optimize the loading priority of learning content on the next launch based on attention metrics, including: After the science popularization content presentation ends, collect user interaction behavior data during this interaction process; Based on the collected user interaction behavior data, a comprehensive calculation is performed using a weighting formula to obtain an indicator that quantifies users' interest in different popular science content. Based on the calculated attention index, the loading priority weight of the corresponding popular science content in the content library is dynamically updated to obtain the updated loading priority weight. Based on the updated loading priority weights and combined with exhibition equipment using standardized interfaces, the optimized content loading order is persistently stored in the database. The standardized exhibition equipment supports flexible replacement of display content and physical form, enabling the booth to adapt to different exhibition themes according to the updated priority, thus achieving persistent application of the exhibition.
8. An AI-powered interactive learning exhibition booth intelligent control system, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The acquisition module is used to detect the arrival of users based on the weight sensor and start the process, identify user group attributes through the camera, and load the appropriate exhibition content mode. The adjustment module is used to collect environmental visual and audio information after the content mode is loaded to obtain initial environmental interference assessment parameters. The camera uses three fixed physical structures on the exhibition stand as monitoring points: the upper left corner of the interactive main screen, the upper right corner of the control panel, and the central marker point on the front edge of the exhibition stand, to obtain a preset marking area. The preset marking area is divided into multiple regular sub-regions. The changes in the imaging characteristics of each sub-region in the visual image are analyzed to obtain the environmental adjustment coefficient. The initial environmental interference assessment parameters are calibrated using the environmental adjustment coefficient to obtain the calibrated environmental interference assessment parameters. The calculation module is used to dynamically adjust the confidence thresholds of gesture trajectory and semantic understanding based on the calibrated environmental interference assessment parameters, so as to obtain the adjusted confidence thresholds of gesture trajectory and semantic understanding. Based on the adjusted confidence thresholds of gesture trajectory and semantic understanding, the module recognizes user gestures or voice commands. Voice commands become effective when the confidence level exceeds the threshold. Gesture commands must simultaneously meet the confidence thresholds and pass the verification by judging the intersection of the gesture trajectory with a preset line segment to confirm the validity of the command, thereby controlling the multimedia device to display popular science content. The processing module is used to analyze user interaction data to obtain attention metrics after the science popularization content display ends; and to optimize the loading priority of the displayed content on the next startup based on the attention metrics.
9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Alarm method and device
CN105488965A
Exhibition booth control method
CN105843099A
Voice interaction method and voice interaction device
CN107316643A
Point-to-point intermittent training video transmission system and method
CN109451257A
Intelligent interaction system and method for exhibition and display
CN111401952A