A method and device for early warning of risk of disease infection
Patent Information
- Application Number
- CN202610527187.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-04-21
AI Technical Summary
[0005]有鉴于此,本申请提供了一种疾病感染风险的预警方法及装置,主要目的在于改善目前现有技术进行呼吸道传染病监测,会出现监测不准确、监测疏漏等情况;单独使用麦克风采集音频会受到教室里说话、桌椅挪动等噪音干扰,导致无法精准判断学生是否出现咳嗽等症状;单独使用摄像头采集视频会受光线和遮挡的影响,导致无法采集到无明显肢体动作的咳嗽行为,无法精准识别到产生咳嗽等症状的学生,进而导致监测的准确度低,无法对呼吸道传染病进行准确预警的技术问题
[0016]借由上述技术方案,本申请提供的一种疾病感染风险的预警方法及装置,与目前现有技术相比,本申请通过基于目标教室的学生座位信息对目标音频进行声源定位并完成座位信息的一致性校验,提升风险学生座位定位的准确性;通过在座位信息符合一致性条件时基于目标视频对风险学生进行感染风险动作识别,实现音视频多模态的感染风险验证,提升识别疾病感染的准确度;通过识别到感染风险动作后确定疾病感染信息并在目标教室进行疾病传染预警,提升教室疾病感染风险预警的有效性。
Smart Images

Figure CN122091264B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of educational technology, and in particular to a method and device for early warning of disease infection risk. Background Technology
[0002] Classrooms are the core teaching areas in educational institutions such as primary and secondary schools and universities, characterized by high population density, relatively fixed seating, and limited air circulation. Respiratory infectious diseases are easily transmitted in classrooms, with early transmission manifesting as external symptoms such as coughing and sneezing.
[0003] Currently, the relevant technologies use microphones to collect audio or cameras to collect video to monitor respiratory infectious diseases. This allows for early warning of respiratory infectious diseases in classrooms once the number of times students exhibit symptoms such as coughing in the audio or video reaches a fixed threshold.
[0004] However, using this method for respiratory infectious disease monitoring can lead to inaccurate monitoring and omissions. Using a microphone alone to collect audio is susceptible to noise interference from talking in the classroom and moving desks and chairs, making it impossible to accurately determine whether students are coughing or have other symptoms. Using a camera alone to collect video is affected by lighting and obstructions, making it impossible to capture coughing behavior without obvious limb movements, and thus impossible to accurately identify students who are coughing or have other symptoms. Consequently, the accuracy of monitoring is low, and it is impossible to provide accurate early warnings for respiratory infectious diseases. Summary of the Invention
[0005] In view of this, this application provides a method and device for early warning of disease infection risk. The main purpose is to improve the existing technology for monitoring respiratory infectious diseases, which suffers from inaccurate monitoring and omissions; using a microphone alone to collect audio is affected by noise such as talking and moving desks and chairs in the classroom, making it impossible to accurately determine whether students have symptoms such as coughing; using a camera alone to collect video is affected by light and obstruction, making it impossible to capture coughing behavior without obvious limb movements, and thus impossible to accurately identify students with coughing symptoms, resulting in low monitoring accuracy and the inability to accurately warn of respiratory infectious diseases.
[0006] Firstly, this application provides a method for early warning of disease infection risk, including: Acquire audio and video data of the target classroom, perform voiceprint recognition on the target audio in the audio and video data, and identify the student at risk of disease infection corresponding to the target audio. The target audio is a sound made through the nasal cavity and / or throat when infected with a disease. Based on the student seating information of the target classroom, the sound source of the target audio is located to obtain the first seating information corresponding to the target audio, and the consistency of the seating information is verified based on the first seating information and the second seating information corresponding to the student at risk. If the first seat information and the second seat information meet the consistency condition, the infection risk action identification is performed on the student at risk based on the target video corresponding to the target audio in the audio and video data. The infection risk action includes facial movements and / or limb movements. In response to the identification of the infection risk action of the student at risk, the disease infection information corresponding to the student at risk is determined, and a disease transmission warning is issued in the target classroom based on the disease infection information.
[0007] Optionally, based on the student seating information of the target classroom, the target audio is localized to obtain the first seating information corresponding to the target audio, and the consistency of the seating information is verified based on the first seating information and the second seating information corresponding to the at-risk student, including: Determine the horizontal azimuth and radial distance information between the sound source location of the target audio and the microphone in the target classroom, and perform sound source localization on the target audio based on the horizontal azimuth and radial distance information to obtain the sound source coordinate information of the target audio; Based on the positioning accuracy of the sound source coordinate information, the sound source range corresponding to the sound source coordinate information is determined, and the first seat within the sound source range is selected from multiple student seats corresponding to the target classroom according to the student seat information, and the first seat information is determined based on the first seat. The second seat information corresponding to the student at risk is determined from the student seat information, and the consistency of the seat information is checked based on the first seat information and the second seat information to obtain the consistency check result corresponding to the first seat information and the second seat information.
[0008] Optionally, before performing sound source localization on the target audio based on the student seating information of the target classroom to obtain the first seating information corresponding to the target audio, and performing a consistency check on the seating information based on the first seating information and the second seating information corresponding to the at-risk student, the method further includes: Based on the lens parameters, image parameters, and center pixel coordinates of the camera in the target classroom, the pixel offset angle of the camera relative to the student's face is determined; Based on the pixel offset angle, the height of the student's face from the ground, the camera's pose parameters, and the camera's global coordinates, the student's face coordinates in the target classroom are determined. The face coordinate information is mapped to the seating layout in the target classroom to obtain the student seating information.
[0009] Optionally, when the first seat information and the second seat information meet the consistency condition, the step of identifying the infection risk actions of the at-risk student based on the target video corresponding to the target audio in the audio and video data includes: If the first seat information and the second seat information meet the consistency condition, the target timestamp of the target audio is determined from the timestamp set corresponding to the audio and video data; Select video data corresponding to the target timestamp from the audio and video data, and determine the selected video data as the target video corresponding to the target audio. Then, determine a set of candidate video frames containing the action scenes of the student at risk from the target video. Based on the candidate video frame set, facial motion feature recognition is performed on the facial movements of the at-risk students, and / or limb motion feature recognition is performed on the limb movements of the at-risk students, in order to identify whether there are infection risk movements in the candidate video frame set that meet the conditions for disease infection movements.
[0010] Optionally, in response to identifying the infection risk action of the at-risk student, determining the disease infection information corresponding to the at-risk student, and issuing a disease transmission warning in the target classroom based on the disease infection information, including: Based on the second seat information and the historical seat information of historically risky students, seat adjacency relationship analysis and seat clustering relationship analysis are performed to obtain the seat clustering area corresponding to the risky students; Based on the clustering scale parameters, clustering density parameters, and contact risk parameters corresponding to the seat clustering area, the disease infection seat information corresponding to the risky student is obtained through fusion analysis. Time sequence analysis and time interval analysis are performed on the target timestamp and the historical target audio timestamp to obtain the time characteristic information corresponding to the risk student; By fusing and analyzing the time trend parameters, cumulative scale parameters, and type parameters corresponding to the time feature information, the disease infection time information corresponding to the at-risk students can be obtained. Based on the disease infection seat information and disease infection time information in the disease infection information, a disease transmission warning is issued in the target classroom.
[0011] Optionally, the step of fusing and analyzing the cluster size parameters, cluster density parameters, and contact risk parameters corresponding to the seat clustering area to obtain the disease infection seat information corresponding to the at-risk students includes: Based on the frequency of occurrence of the target audio and historical target audio in the seat clustering area and the frequency of occurrence of the target audio and historical target audio in the target classroom, the clustering scale parameter of the seat clustering area is determined; The clustering density parameter is determined based on the frequency of occurrence of the target audio and historical target audio in the seat clustering area and the area of the seat clustering area. The contact risk parameters are determined based on the number of at-risk students in the seating area and the number of at-risk students in the target classroom. The first importance parameter of the cluster size parameter, the cluster density parameter, and the contact risk parameter is determined, and the cluster size parameter, the cluster density parameter, and the contact risk parameter are fused based on the first importance parameter to obtain the disease infection seat information.
[0012] Optionally, the time trend parameters, cumulative scale parameters, and type parameters corresponding to the time feature information are fused and analyzed to obtain the disease infection time information corresponding to the at-risk students, including: The time trend parameter is determined based on the changing trend of the time sequence in the time feature information; The cumulative scale parameter is determined based on the number of times the target audio appears and the number of students in the target classroom; The proportion of consecutive timestamps in the target timestamp is determined as the type parameter; A second importance parameter is determined for the time trend parameter, the cumulative scale parameter, and the type parameter, and the time trend parameter, the cumulative scale parameter, and the type parameter are fused based on the second importance parameter to obtain the disease infection time information.
[0013] Optionally, the step of issuing a disease transmission warning in the target classroom based on the disease infection seat information and the disease infection time information in the disease infection information includes: The disease infection seat information, the disease infection time information, and the air circulation parameters of the target classroom are fused together to obtain the disease transmission risk data of the target classroom; The number of students at risk and students with historical risk, the number of areas where seats are clustered, and the disease transmission risk data are merged to obtain comprehensive risk data. Risk level analysis is performed on the comprehensive risk data to obtain the disease transmission risk level corresponding to the comprehensive risk data; Based on the disease transmission risk level, a warning message corresponding to the disease transmission risk level is generated, and a disease transmission warning is issued in the target classroom according to the warning message. The warning message includes the second seat information of the at-risk student, the target timestamp of the target audio, the disease infection seat information, the disease infection time information, the disease transmission risk data, the disease transmission risk level, and disease transmission prevention suggestions.
[0014] Optionally, the step of acquiring the audio and video data of the target classroom, performing voiceprint recognition on the target audio in the audio and video data, and determining the student at risk of disease infection corresponding to the target audio includes: Audio data within the target frequency range is selected from the audio data, and the selected audio data is determined as the target audio. Determine the target voiceprint information corresponding to the target audio, and match the target voiceprint information with the voiceprint information of multiple students to obtain voiceprint matching degree data; If the voiceprint matching data meets the voiceprint matching conditions, the student corresponding to the target voiceprint information is identified as a student at risk of disease infection.
[0015] Secondly, this application provides an early warning device for the risk of disease infection, comprising: The acquisition module is configured to acquire audio and video data of the target classroom, perform voiceprint recognition on the target audio in the audio and video data, and determine the student at risk of disease infection corresponding to the target audio, wherein the target audio is a sound emitted through the nasal cavity and / or throat when infected with a disease; The verification module is configured to locate the sound source of the target audio based on the student seating information of the target classroom, obtain the first seating information corresponding to the target audio, and perform a consistency verification of the seating information based on the first seating information and the second seating information corresponding to the student at risk. The identification module is configured to identify infection risk actions of the at-risk student based on the target video corresponding to the target audio in the audio and video data, when the first seat information and the second seat information meet the consistency condition. The infection risk actions include facial movements and / or limb movements. The early warning module is configured to respond to the identification of the infection risk action of the at-risk student, determine the disease infection information corresponding to the at-risk student, and issue a disease transmission warning in the target classroom based on the disease infection information.
[0016] By employing the above technical solutions, this application provides a method and apparatus for early warning of disease infection risk. Compared with existing technologies, this application improves the accuracy of locating the seats of at-risk students by locating the sound source of target audio based on the student seating information of the target classroom and completing the consistency verification of the seating information; it also improves the accuracy of disease infection identification by recognizing infection risk actions of at-risk students based on the target video when the seating information meets the consistency conditions; and it improves the effectiveness of classroom disease infection risk early warning by determining disease infection information after recognizing infection risk actions and issuing disease transmission early warning in the target classroom. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a method for early warning of disease infection risk provided in an embodiment of this application is shown. Figure 2 A flowchart illustrating a method for early warning of disease infection risk provided in an embodiment of this application is shown. Figure 3 A schematic diagram of a three-dimensional coordinate system in a classroom provided in an embodiment of this application is shown; Figure 4 A schematic diagram of the structure of a disease infection risk early warning device provided in an embodiment of this application is shown; Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0020] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0021] To improve existing technologies for monitoring respiratory infectious diseases, issues such as inaccurate monitoring and omissions arise. Using microphones alone for audio recording is susceptible to noise interference from classroom conversations and furniture movement, making it difficult to accurately determine if students are coughing or exhibiting other symptoms. Using cameras alone for video recording is affected by lighting and obstructions, preventing the capture of coughing behaviors without obvious limb movements and hindering the accurate identification of students exhibiting coughing symptoms. These issues lead to low monitoring accuracy and an inability to provide accurate early warnings for respiratory infectious diseases. This embodiment provides a method for early warning of disease infection risk, such as... Figure 1 As shown, the method includes: Step 101: Obtain the audio and video data of the target classroom, perform voiceprint recognition on the target audio in the audio and video data, and identify the student at risk of disease infection corresponding to the target audio.
[0022] The target audio is the sound emitted through the nasal cavity and / or throat in the case of an infectious disease.
[0023] In this embodiment, the target classroom can be a place for conducting teaching activities, such as student learning and teacher instruction. For example, the target classroom in this embodiment can specifically include classrooms in primary and secondary schools, universities, and various training institutions, and can also be extended to places with dense populations and relatively fixed seating, such as offices and conference rooms.
[0024] In the embodiments of this application, audio and video data can be a collection of sound and image data collected by audio acquisition devices and video acquisition devices. For example, the audio and video data in the embodiments of this application may specifically include audio streams collected by multiple array microphones in a classroom and video streams collected by a high-definition camera. The collection of audio and video data can be done during classroom teaching; during breaks, due to students' changing positions and noisy audio and video environments, audio and video data are not collected.
[0025] In this embodiment of the application, the target audio can be audio emitted by a student related to disease infection. For example, the target audio in this embodiment of the application can specifically be sounds such as coughing, sneezing, and blowing one's nose emitted by the human body through the nasal cavity and / or throat when infected with a respiratory infectious disease.
[0026] In this embodiment of the application, voiceprint recognition can involve extracting voiceprint features from audio and matching them with preset features. Voiceprint recognition can be used to identify students who emit target audio. For example, the voiceprint recognition in this embodiment may specifically include processing steps such as feature extraction, feature normalization, and cosine similarity matching of the audio.
[0027] In this embodiment of the application, a student at risk can be a student exhibiting vocal behaviors related to disease infection, and such students can be considered key targets for disease infection monitoring and early warning. For example, in this embodiment, a student at risk can specifically be a student whose voiceprint information matches the voiceprint information of the target audio file according to preset conditions.
[0028] In this embodiment, audio and video data are collected using audio and video acquisition devices deployed in the target classroom. The audio acquisition device can be a mainstream low-cost multi-array microphone used on campus, deployed in the center of the classroom ceiling. The video acquisition device can be a high-definition camera, deployed at the front, middle, left and right sides of the classroom to capture students' faces without obstruction. The audio data in the collected audio and video data is preliminarily processed to filter out target audio that matches the vocal characteristics related to disease infection. Voiceprint recognition is performed on the target audio to extract the voiceprint features and match them with the voiceprint information of students in the classroom. Students whose voiceprint matching results meet preset conditions are identified as students at risk of disease infection.
[0029] Step 102: Based on the student seating information of the target classroom, perform sound source localization on the target audio to obtain the first seat information corresponding to the target audio, and perform consistency verification of the seat information based on the first seat information and the second seat information corresponding to the risky student.
[0030] In this embodiment of the application, student seating information can be information representing the spatial location and correspondence of student seats in the classroom. This information can be used for matching sound source localization results and verifying the consistency of seating information. For example, the student seating information in this embodiment may specifically include the three-dimensional spatial coordinates of each seat in the classroom, the seat number, and the student's identity information corresponding to that seat.
[0031] In the embodiments of this application, sound source localization can be achieved by determining the location of the sound source based on the characteristics of the audio signal, and sound source localization can be used to determine the location corresponding to the target audio. For example, sound source localization in the embodiments of this application may specifically include determining the sound source location using the hardware interface (Application Programming Interface, API) of a multi-array microphone combined with a spatial coordinate transformation algorithm.
[0032] In this embodiment of the application, the first seat information may be seat information corresponding to the target audio obtained through sound source localization. For example, the first seat information in this embodiment of the application may specifically include the seat spatial coordinates, seat number, and sound source range of the seat obtained through sound source localization.
[0033] In this embodiment, the second seat information may be the actual seat information of the at-risk student in the classroom. For example, the second seat information in this embodiment may specifically include the spatial coordinates and seat number of the at-risk student obtained through face positioning and seat layout mapping.
[0034] In this embodiment, the consistency verification of seat information can be a process of comparing first seat information and second seat information and determining whether they match. For example, the consistency verification of seat information in this embodiment can specifically include a verification process of determining whether the seat corresponding to the second seat information is within the sound source range corresponding to the first seat information.
[0035] Step 103: If the first seat information and the second seat information meet the consistency condition, identify the infection risk actions of students based on the target video corresponding to the target audio in the audio and video data.
[0036] Among these, actions that pose a risk of infection include facial movements and / or limb movements.
[0037] In this embodiment of the application, the target video can be video data that matches the temporal characteristics of the target audio, and the target video can be used to identify infection risk actions of at-risk students. For example, in this embodiment of the application, the target video may specifically include a sequence of consecutive video frames aligned with the millisecond-level timestamps of the target audio.
[0038] In this embodiment, the infection risk action can be a related action performed when a target audio message is emitted due to an infectious disease. This infection risk action can serve as a visual feature to further verify the student's disease infection risk. For example, the infection risk action in this embodiment may specifically include facial and / or limb movements related to respiratory infectious diseases.
[0039] Optionally, the infection risk actions in this application embodiment may specifically include only facial actions related to respiratory infectious diseases, only limb actions related to respiratory infectious diseases, or both facial actions and limb actions related to respiratory infectious diseases.
[0040] In this embodiment, facial movements can be related to the facial movements made by a student at risk of contracting a disease when emitting the target audio. For example, facial movements in this embodiment may include, but are not limited to, actions such as opening the mouth, covering the mouth, frowning, tilting the head forward, and lowering the head when coughing.
[0041] In this embodiment of the application, body movements can be related to the body movements made by a student at risk of contracting an illness when emitting the target audio. Body movements can serve as another visual feature for identifying infection risk. For example, body movements in this embodiment of the application may include, but are not limited to, shoulder tremors, chest tremors, and hand movements toward the mouth during coughing.
[0042] In this embodiment of the application, it is determined whether the consistency verification results of the first seat information and the second seat information meet the preset consistency conditions. If they do not meet the conditions, the operation is terminated. If they do meet the conditions, the target video corresponding to the target audio is selected from the collected audio and video data according to the timestamp matching principle. The screen area of the student at risk is located in the target video. Action recognition is performed on the student at risk in the screen area. Feature extraction and analysis are performed on the facial and / or limb movements of the student at risk to determine whether the student at risk has any infection risk actions.
[0043] Step 104: In response to the identification of an infected student at risk, determine the disease infection information corresponding to the student at risk, and issue a disease transmission warning in the target classroom based on the disease infection information.
[0044] In this embodiment of the application, disease infection information can be a collection of information related to disease infection among at-risk students, and this information can be used to assess the risk level of disease transmission within the target classroom. For example, the disease infection information in this embodiment may specifically include information characterizing the spatiotemporal aggregation features of disease infection, such as disease infection seat information and disease infection time information.
[0045] In this embodiment, disease transmission early warning can be achieved by assessing the risk of disease transmission based on disease infection information and triggering corresponding alerts. Disease transmission early warning can be used to achieve early prevention and intervention of respiratory infectious diseases. For example, the disease transmission early warning in this embodiment may specifically include risk level determination, early warning information generation, early warning information push, and visualization result output.
[0046] In this embodiment of the application, if no infection risk action of a student at risk is identified, the operation of this step can be terminated. If an infection risk action of a student at risk is identified, a comprehensive analysis can be performed based on multi-dimensional data such as the student's seat information and the time information of the target audio to determine the disease infection information corresponding to the student at risk. Based on the disease infection information, the risk of disease transmission in the target classroom is quantitatively analyzed and graded. Based on the risk grade determination result, a disease transmission warning is issued in the target classroom.
[0047] Compared with existing technologies, this embodiment improves the accuracy of locating the seats of at-risk students by locating the sound source of the target audio based on the student seating information of the target classroom and completing the consistency verification of the seating information; it also improves the accuracy of monitoring disease infection by recognizing infection risk actions of at-risk students based on the target video when the seating information meets the consistency conditions; and it improves the effectiveness of classroom disease infection risk warning by determining disease infection information after recognizing infection risk actions and issuing disease transmission warnings in the target classroom.
[0048] As an optional approach, when performing the task of "localizing the sound source of the target audio based on the student seating information of the target classroom, obtaining the first seat information corresponding to the target audio, and performing a consistency check of the seating information based on the first seat information and the second seat information corresponding to the at-risk students," the following methods can be used, but are not limited to: Figure 2 As shown, the method includes: Step 201: Determine the horizontal azimuth and radial distance information between the sound source location of the target audio and the microphone in the target classroom, and locate the sound source of the target audio based on the horizontal azimuth and radial distance information to obtain the sound source coordinate information of the target audio.
[0049] In this embodiment, the horizontal azimuth information can be the angle information of the sound source relative to the horizontal orientation of the microphone. For example, in this embodiment, the horizontal azimuth information can specifically be the horizontal azimuth angle directly output by the multi-array microphone API. The angle is based on the Y-axis of the classroom directly in front of the microphone, with a rightward deviation being positive.
[0050] In this embodiment, the radial distance information can be the length of the straight-line distance from the sound source to the microphone. For example, in this embodiment, the radial distance information can specifically be the straight-line distance *r* from the sound source to the microphone output by the multi-array microphone API.
[0051] In this embodiment of the application, the sound source coordinate information can be the coordinate information of the global physical location of the sound source within the target classroom. The sound source coordinate information can be used to determine the seating range corresponding to the sound source. For example, the sound source coordinate information in this embodiment can specifically be the sound source plane coordinates (Xs, Ys) in a three-dimensional coordinate system within the classroom, as illustrated in the diagram below. Figure 3 As shown, Figure 3 Taking a camera as an example, the X-axis, Y-axis, and Z-axis are shown as dashed lines.
[0052] In this embodiment of the application, the horizontal azimuth and radial distance between the sound source location and the microphones can be obtained through the API of the multi-array microphones deployed in the target classroom; based on the microphone's coordinates (Xm, Ym) in the target classroom, the X-axis and Y-axis offsets of the sound source relative to the microphones are calculated using a spatial transformation algorithm from polar coordinates to planar coordinates. The X-axis offset of the sound source relative to the microphones is... The calculation formula is shown in Formula 1, where r can represent the radial distance from the sound source to the microphone. It can represent the horizontal azimuth angle from the sound source to the microphone; the Y-axis offset of the sound source relative to the microphone. The calculation formula can be shown in Formula 2, where r can represent the radial distance from the sound source to the microphone. This can represent the horizontal azimuth angle from the sound source to the microphone; the planar coordinates (Xs, Ys) of the sound source are calculated using the offset. The formula for calculating Xs is shown in Formula 3, where Xm can represent the X-axis coordinate of the microphone within the classroom. Ys can represent the X-axis offset of the sound source relative to the microphone; the formula for calculating Ys is shown in Formula 4, where Ym can represent the Y-axis coordinate of the microphone within the classroom. It can represent the Y-axis offset of the sound source relative to the microphone.
[0053] (Formula 1) (Formula 2) (Formula 3) (Formula 4) Step 202: Based on the positioning accuracy of the sound source coordinate information, determine the sound source range corresponding to the sound source coordinate information, and select the first seat within the sound source range from multiple student seats corresponding to the target classroom according to the student seat information, and determine the first seat information based on the first seat.
[0054] In this embodiment, the positioning accuracy can be a parameter representing the microphone sound source positioning error range, and the positioning accuracy can be used to determine the actual spatial coverage area of the sound source. For example, the positioning accuracy in this embodiment may specifically include the horizontal azimuth measurement error of the multi-array microphone. Radial distance measurement error And the overall planar positioning accuracy calculated from both.
[0055] In this embodiment, the sound source range can be a spatial area where the sound source may exist, determined based on the sound source coordinate information and positioning accuracy. The sound source range can be used to filter the seat corresponding to the target audio. For example, in this embodiment, the sound source range can be a circular spatial area centered on the sound source coordinate information and with the planar comprehensive positioning accuracy as the radius, or it can be an area enclosed by the X-axis and Y-axis direction error intervals determined based on the coordinate increment error.
[0056] In this embodiment of the application, the first seat can be a student seat in the target classroom within the sound source range, and the first seat can serve as the basis for determining the first seat information. For example, in this embodiment of the application, the first seat can specifically be all student seats within the sound source range. The sound source range at a close distance of 1-3 meters in the front / middle rows of the classroom can cover the seats of 2-4 students, and the sound source range at a medium-to-long distance of 5-8 meters in the back rows can cover the seats of 4-10 students.
[0057] For the embodiments of this application, the positioning accuracy of the sound source coordinate information is calculated based on the positioning accuracy parameters of the multi-array microphone. The formula for calculating the coordinate increment error of the X-axis can be shown in Formula 5, where, This can represent the X-axis offset of the sound source relative to the microphone, and r can represent the radial distance from the sound source to the microphone. It can represent the horizontal azimuth angle from the sound source to the microphone. It can represent the horizontal azimuth measurement error of a multi-array microphone. This can represent the radial distance measurement error of a multi-array microphone; the formula for calculating the coordinate increment error of the Y-axis can be shown in Formula 6, where, This can represent the Y-axis offset of the sound source relative to the microphone, and r can represent the radial distance from the sound source to the microphone. It can represent the horizontal azimuth angle from the sound source to the microphone. It can represent the horizontal azimuth measurement error of a multi-array microphone. The radial distance measurement error of a multi-array microphone can be represented; the planar position error range of the sound source on the X-axis of the classroom coordinate system can be represented as... The error range of the sound source's planar position on the Y-axis of the classroom coordinate system can be expressed as: The formula for calculating the planar integrated positioning accuracy R can be shown in Formula 7, where, This can represent the incremental error of the X-axis coordinate. It can represent the incremental error of the Y-axis coordinate.
[0058] (Formula 5) (Formula 6) (Formula 7) In this embodiment of the application, the sound source range corresponding to the sound source coordinate information is determined based on the calculated positioning accuracy; based on the student seat information of the target classroom, the seat within the sound source range is selected from multiple student seats in the classroom and determined as the first seat; the spatial coordinates, seat number and other related information of the first seat are determined as the first seat information corresponding to the target audio.
[0059] Step 203: Determine the second seat information corresponding to the risky student from the student seat information, and perform a consistency check on the seat information based on the first seat information and the second seat information to obtain the consistency check result corresponding to the first seat information and the second seat information.
[0060] In this embodiment, the second seat information corresponding to the at-risk student can be extracted from the student seating information of the target classroom. The second seat information includes the spatial coordinates, seat number, and other relevant information of the at-risk student's actual seat in the classroom. The sound source range, seat coordinates, and other information corresponding to the first seat information are compared with the seat coordinates, seat number, and other information corresponding to the second seat information to obtain a consistency verification result. If the seat corresponding to the second seat information is within the sound source range corresponding to the first seat information, the verification result is determined to meet the consistency condition. If the seat corresponding to the second seat information is not within the sound source range corresponding to the first seat information, the verification result is determined to not meet the consistency condition.
[0061] Optionally, before performing the step of "localizing the target audio source based on the student seating information in the target classroom to obtain the first seat information corresponding to the target audio, and performing a consistency check on the seat information based on the first seat information and the second seat information corresponding to the at-risk student," the following method may be used, but is not limited to: determining the pixel offset angle of the camera relative to the student's face based on the lens parameters, image parameters, and center pixel coordinates of the camera in the target classroom; determining the student's face coordinate information in the target classroom based on the pixel offset angle, the height of the student's face from the ground, the camera's pose parameters, and the camera's global coordinates; and mapping the face coordinate information to the seating layout in the target classroom to obtain the student's seating information.
[0062] In this embodiment of the application, the lens parameters can be the optical characteristic parameters of the camera, and the lens parameters can be used to calculate the pixel offset angle of the camera relative to the student's face. For example, the lens parameters in this embodiment of the application may specifically include parameters such as the horizontal field of view (HFOV) and the vertical field of view (VFOV) of the camera.
[0063] In the embodiments of this application, image parameters can be pixel feature parameters of the image captured by the camera, and image parameters can be used to calculate pixel deviation angles. For example, the image parameters in the embodiments of this application may specifically include pixel parameters such as the width W and height H of the image captured by the camera.
[0064] In this embodiment of the application, the center pixel coordinates can be the pixel position information of the student's face center in the image captured by the camera. The center pixel coordinates can be used to determine the specific position of the face in the image. For example, in this embodiment of the application, the center pixel coordinates can specifically be the pixel coordinates (u,v) of the student's face center in the image.
[0065] In this embodiment, the pixel offset angle can be the offset angle of the camera's optical axis relative to the student's face. The pixel offset angle can be used to calculate the unit direction vector of the camera pointing towards the face. For example, the pixel offset angle in this embodiment may specifically include a horizontal offset angle. and vertical deflection angle .
[0066] In this embodiment, the pose parameter can be the camera's mounting angle parameter, and the pose parameter can be used to calculate the global physical coordinates of the student's face. For example, the pose parameter in this embodiment may specifically include the horizontal angle of the camera. Pitch angle Parameters such as these.
[0067] In this embodiment, the global coordinates can be the three-dimensional spatial position information of the camera within the target classroom, and can be used as the reference coordinates for calculating the student's face coordinates. For example, in this embodiment, the global coordinates can specifically be the three-dimensional spatial coordinates (Xc, Yc, Zc) of the camera within the classroom.
[0068] In this embodiment of the application, the face coordinate information can be the global physical location information of the student's face within the target classroom, and the face coordinate information can be used to map and match with the classroom seating layout. For example, in this embodiment of the application, the face coordinate information can specifically be the two-dimensional planar coordinates (Xp, Yp) of the student's face within the classroom.
[0069] In this embodiment, the seating layout can be information about the arrangement and coordinate correspondence of seats in the target classroom. The seating layout can be used to map face coordinate information to student seating information. For example, the seating layout in this embodiment can specifically include the global physical coordinates of each seat in the classroom, the seating arrangement, row and column distribution, etc.
[0070] In this embodiment of the application, based on lens parameters such as the horizontal and vertical field of view of the camera in the target classroom, and image parameters such as the width and height of the acquired image, combined with the pixel coordinates (u,v) of the student's face center, the pixel deviation angle of the camera relative to the student's face is calculated, including the horizontal deviation angle. The calculation formula is shown in Formula 8, where u represents the X-axis pixel coordinates of the student's face center in the image, W represents the width of the image captured by the camera, HFOV represents the horizontal field of view of the camera, and the vertical tilt angle represents the vertical angle. The calculation formula can be shown in Formula 9, where v can represent the Y-axis pixel coordinate of the student's face center in the image, H can represent the height of the image captured by the camera, and VFOV can represent the vertical field of view of the camera. (Formula 8) (Formula 9) In this embodiment, based on the pixel offset angle, combined with the height Zp of the student's face from the ground and the horizontal angle of the camera, Pitch angle Given the pose parameters and the camera's 3D global coordinates (Xc, Yc, Zc), calculate the unit direction vector of the camera pointing towards the student's face, and the unit direction vector of the camera pointing towards the student's face along the X-axis. The calculation formula can be shown in Formula 10, where, This can represent the horizontal tilt angle that indicates the angle at which a camera pixel deviates from its position. It can represent the vertical tilt angle of a camera pixel's deviation angle. It can indicate the horizontal angle at which the camera is installed. This can represent the vertical angle at which the camera is mounted; and the unit Y-axis direction vector of the camera pointing towards the student's face. The calculation formula can be shown in Formula 11, where, This can represent the horizontal tilt angle that indicates the angle at which a camera pixel deviates from its position. It can represent the vertical tilt angle of a camera pixel's deviation angle. It can indicate the horizontal angle at which the camera is installed. This can represent the vertical angle at which the camera is mounted; and the unit Z-axis direction vector of the camera pointing towards the student's face. The calculation formula can be shown in Formula Twelve, where, This can represent the horizontal tilt angle that indicates the angle at which a camera pixel deviates from its position. It can represent the vertical tilt angle of a camera pixel's deviation angle. It can indicate the horizontal angle at which the camera is installed. It can indicate the vertical angle at which the camera is installed.
[0071] (Formula 10) (Formula Eleven) (Formula 12) In this embodiment, the formula for calculating the camera ray scaling factor k is as shown in Formula Thirteen, where Zp represents the height of the student's face from the ground, and Zc represents the Z-axis coordinate of the camera within the classroom. The Z-axis unit direction vector of the camera pointing at the student's face can be represented. The formula for calculating the X-axis coordinate Xp of the student's face within the classroom using the ray scaling factor and the unit direction vector is shown in Formula Fourteen, where Xc represents the X-axis coordinate of the camera within the classroom, and k represents the ray scaling factor of the camera. The X-axis unit direction vector of the camera pointing at the student's face can be represented; the formula for calculating the Y-axis coordinate Yp of the student's face in the classroom can be shown in Formula 15, where Yc can represent the Y-axis coordinate of the camera in the classroom, and k can represent the ray scaling factor of the camera. The Y-axis unit direction vector of the camera pointing at the student's face can be represented, and (Xp, Yp) is determined as the face coordinate information. The face coordinate information is mapped and matched with the seating layout in the target classroom. Based on the global physical coordinates of each seat in the seating layout, the face coordinate information is matched to the corresponding student seat, and then the relevant information of the seat is extracted to obtain the corresponding student seat information.
[0072] (Formula Thirteen) (Formula Fourteen) (Formula 15) As an optional approach, when performing the task of "identifying infection risk actions of at-risk students based on the target video corresponding to the target audio in the audio-video data, provided that the first and second seat information meet the consistency condition," the following method can be used, but is not limited to: determining the target timestamp of the target audio from the set of timestamps corresponding to the audio-video data, provided that the first and second seat information meet the consistency condition; selecting video data corresponding to the target timestamp from the audio-video data, and determining the selected video data as the target video corresponding to the target audio, and determining a set of candidate video frames containing the action scenes of at-risk students from the target video; and, based on the set of candidate video frames, performing facial action feature recognition on the facial movements of at-risk students, and / or performing limb action feature recognition on the limb movements of at-risk students, to identify whether there are infection risk actions in the set of candidate video frames that meet the conditions for disease infection actions.
[0073] In the embodiments of this application, the timestamp set can be the time set corresponding to each frame of audio and video data. For example, in the embodiments of this application, the timestamp set can specifically be a set of millisecond-level time stamps synchronized using the NTP network time protocol, with an audio and video synchronization error of less than or equal to 10ms.
[0074] In this embodiment of the application, the target timestamp can be time information corresponding to the target audio, and the target timestamp can be used to filter the video data corresponding to the target audio. For example, in this embodiment of the application, the target timestamp can specifically be a millisecond-level time range consisting of the start timestamp and end timestamp of the target audio.
[0075] In this embodiment, the candidate video frame set can be a series of video frames from the target video that contain images of the actions of students at risk. The candidate video frame set can be used for feature recognition of actions that pose a risk of infection. For example, in this embodiment, the candidate video frame set can specifically be a series of video frames extracted from the target video at a frame rate greater than or equal to 25 fps that contain the upper body region of students at risk.
[0076] In this embodiment, the disease infection action condition can be a feature standard for judging the risk of infection. The disease infection action condition can be used to screen facial and / or limb movements that meet the requirements. For example, the disease infection action condition in this embodiment may specifically include mouth opening and closing degree, shoulder tremor pixel displacement variance, etc.
[0077] In this embodiment of the application, when the consistency verification results of the first seat information and the second seat information meet the consistency conditions, the start time stamp and end time stamp of the target audio are extracted from the time-synchronized timestamp set corresponding to the audio and video data, and determined as the target timestamp; based on the target timestamp, the corresponding video data is selected from the audio and video data, determined as the target video corresponding to the target audio, and the screen area of the student at risk is located from the target video, and continuous video frames containing the action screen of the student at risk are extracted to form a candidate video frame set.
[0078] In this embodiment of the application, the candidate video frame set is preprocessed, including operations such as grayscale conversion and Gaussian filtering for noise reduction. Based on preset disease infection action conditions, facial action feature recognition is performed on the facial movements of at-risk students, and / or limb action feature recognition is performed on the limb movements of at-risk students. For example, facial action recognition can extract key facial points such as the mouth, nose, and eye sockets to analyze relevant features, while limb action recognition can extract key limb points such as the shoulder, chest, and hand to analyze relevant features.
[0079] Optionally, inter-frame differences can be calculated by analyzing the coordinates of key points such as the mouth, nose, eye sockets, shoulders, chest, and hands in the candidate video frame set of the target video. The inter-frame difference calculation of key point coordinates for facial and limb motion features can include head posture features, mouth opening and closing features, mouth-covering action features, and shoulder / chest trembling features.
[0080] For example, when the target audio is a cough sound, head posture features can be calculated by measuring the pitch or forward tilt angle of the head key points relative to the midpoint of the shoulder; if the change in pitch or forward tilt angle is greater than or equal to 15 degrees for more than 3 consecutive frames, then the head motion features can be determined to match the cough.
[0081] For example, when the target audio is a cough sound, the mouth opening and closing feature can be determined by calculating the Euclidean distance between the key points of the upper and lower lips. If the mouth opening and closing feature is greater than or equal to 0.3 (open mouth) for more than 2 consecutive frames, or if the mouth opening and closing feature fluctuates rapidly between 0.1 and 0.3 (rapid opening and closing of the mouth during coughing), then the mouth action feature can be determined to match the cough.
[0082] For example, when the target audio is a coughing sound, the mouth-covering action feature can be determined by calculating the Euclidean distance between the key points of any fingertips of the left and right hands and the center point of the mouth; if the Euclidean distance between the key points of any fingertips of the left and right hands and the center point of the mouth is less than or equal to 10 pixels for more than 2 consecutive frames, then the mouth-covering action feature can be determined to match the cough.
[0083] For example, when the target audio is a cough sound, the shoulder / chest tremor feature can be determined by calculating the pixel displacement variance of the shoulder key points in consecutive frames; if the pixel displacement variance is greater than or equal to 5, the shoulder / chest tremor feature can be determined to match the cough.
[0084] Optionally, if both the mouth opening and closing features and the shoulder / chest tremor features match a cough, the student at risk can be identified as having a coughing action; if the mouth covering or head movement features match a cough and the shoulder / chest tremor features match a cough, the student at risk can be identified as having a coughing action; if none of the above conditions are met, the student at risk can be identified as not having a coughing action.
[0085] Optionally, embodiments of this application may, based on a set of candidate video frames, perform facial motion feature recognition only on the facial movements of at-risk students, or perform limb motion feature recognition only on the limb movements of at-risk students, or perform both facial motion feature recognition and limb motion feature recognition on the facial movements of at-risk students.
[0086] As an optional approach, when performing the action of "responding to the identification of infection risk of at-risk students, determining the disease infection information corresponding to the at-risk students, and issuing a disease transmission warning in the target classroom based on the disease infection information," the following methods can be used, but are not limited to: analyzing seat adjacency and seat clustering relationships based on the second seat information and the historical seat information of historical at-risk students to obtain the seat clustering area corresponding to the at-risk students; performing a fusion analysis based on the clustering scale parameters, clustering density parameters, and contact risk parameters corresponding to the seat clustering area to obtain the disease infection seat information corresponding to the at-risk students; performing time sequence and time interval analysis on the target timestamp and the historical target timestamp of the historical target audio to obtain the time characteristic information corresponding to the at-risk students; performing a fusion analysis based on the time trend parameters, cumulative scale parameters, and type parameters corresponding to the time characteristic information to obtain the disease infection time information corresponding to the at-risk students; and issuing a disease transmission warning in the target classroom based on the disease infection seat information and disease infection time information in the disease infection information.
[0087] In this embodiment, "historically risky students" can be students previously identified as having a risk of disease infection. These students can be used to analyze seating clustering characteristics related to disease infection. For example, in this embodiment, "historically risky students" can specifically be students in the classroom identified as having a risk of disease infection prior to this identification, through voiceprint recognition and motion recognition.
[0088] In this embodiment of the application, historical seat information may be seat-related information corresponding to students with historical risks. This historical seat information can be used for seat adjacency and clustering analysis. For example, in this embodiment, the historical seat information may specifically include the spatial coordinates and seat number of the students with historical risks.
[0089] In this embodiment, seat adjacency analysis can be a process of analyzing the seat adjacency of high-risk students and historically high-risk students. Seat adjacency analysis can be used to analyze the transmission pathways of diseases. For example, the seat adjacency analysis in this embodiment can specifically include the process of analyzing the front-back and left-right adjacent distribution of seats.
[0090] In this embodiment of the application, seat clustering analysis can be a process of analyzing the spatial clustering of multiple high-risk student seats. Seat clustering analysis can be used to determine seat clustering areas associated with disease infection. For example, the seat clustering analysis in this embodiment can specifically employ a density clustering algorithm to perform spatial clustering analysis on high-risk seats.
[0091] In this embodiment, the seat clustering area can be an area within the target classroom where high-risk student seats are spatially clustered. This seat clustering area can be used to assess the spatial transmission risk of disease infection. For example, in this embodiment, the seat clustering area can specifically be a cluster of cough events identified using the DBSCAN density clustering algorithm.
[0092] In this embodiment of the application, the disease-infected seating information can be disease-infected spatial characteristic information obtained by fusing spatial clustering parameters. This disease-infected seating information can be used to assess the spatial transmission risk of the disease. For example, in this embodiment of the application, the disease-infected seating information can specifically be the Classroom Spatial Clustering Index (CSI).
[0093] In this embodiment, the historical target audio can be previously identified vocalizations related to disease infection, which can be used to analyze the temporal evolution characteristics of disease infection. For example, in this embodiment, the historical target audio can specifically be audio such as coughing or sneezing sounds filtered from audio and video data prior to this identification.
[0094] In this embodiment of the application, the historical timestamp can be time stamp information corresponding to the historical target audio, and the historical timestamp can be used for time sequence and time interval analysis. For example, in this embodiment of the application, the historical timestamp can specifically be the start and end timestamps of the sound transmission corresponding to the historical target audio.
[0095] In this embodiment of the application, time sequence analysis can be a process of analyzing the temporal order of the target audio and historical target audio. Time sequence analysis can be used to extract the temporal evolution pattern of disease infection. For example, in this embodiment of the application, time sequence analysis can specifically be a process of statistically analyzing the order of cough events according to time windows.
[0096] In this embodiment of the application, time interval analysis can be a process of analyzing the time interval between the target audio and historical target audio. For example, in this embodiment, time interval analysis can specifically be a process of statistically analyzing the distribution of time intervals between consecutive coughing events.
[0097] In the embodiments of this application, the time feature information can be disease infection time-related feature information obtained through time analysis. The time feature information can be used to quantify the temporal transmission characteristics of disease infection. For example, the time feature information in the embodiments of this application may specifically include the time trend, frequency of occurrence, and type proportion of cough events.
[0098] In this embodiment, the disease infection time information can be disease infection time feature information obtained by fusing time feature parameters. This disease infection time information can be used to assess the risk of disease transmission over time. For example, the disease infection time information in this embodiment may specifically include the Temporal Spread Index (TSI).
[0099] In this embodiment of the application, based on the second seat information of students at risk and the historical seat information of students with historical risk, spatial clustering algorithms, including but not limited to the DBSCAN density clustering algorithm, are used to analyze the seat adjacency relationship and seat clustering relationship of seats in the classroom. The neighborhood radius of the DBSCAN algorithm is used as an example. The value can be set to 1.2m, and the minimum number of points (MinPts) can be set to 3. Clustering is used to identify the seat clustering areas corresponding to high-risk students in the target classroom. Based on the clustering scale parameter, clustering density parameter, and contact risk parameter corresponding to the seat clustering area, a weighted fusion analysis is performed to obtain the disease infection seat information corresponding to the high-risk students.
[0100] In this embodiment of the application, the target timestamp of the target audio and the historical timestamp of the historical target audio are analyzed by time sequence and time interval. The data can be binned according to time windows of 0.5 calendar days, 1 calendar day, and 2 calendar days to extract the time feature information corresponding to the students at risk. Based on the time trend parameters, cumulative scale parameters, and type parameters corresponding to the time feature information, a weighted fusion analysis is performed to obtain the disease infection time information corresponding to the students at risk. The disease infection seat information and disease infection time information are used as the core content of disease infection information. Based on this disease infection information, the risk of disease transmission in the target classroom is quantitatively assessed, and then the corresponding disease transmission early warning operation is executed.
[0101] As an optional approach, when performing the "fusion analysis of the cluster size parameter, cluster density parameter, and contact risk parameter corresponding to the seat cluster area to obtain the disease infection seat information corresponding to the at-risk students", the following methods can also be used, but are not limited to: determining the cluster size parameter of the seat cluster area based on the occurrence frequency of the target audio and historical target audio in the seat cluster area and the occurrence frequency of the target audio and historical target audio in the target classroom; determining the cluster density parameter based on the occurrence frequency of the target audio and historical target audio in the seat cluster area and the area of the seat cluster area; determining the contact risk parameter based on the number of at-risk students in the seat cluster area and the number of at-risk students in the target classroom; determining the first importance parameter of the cluster size parameter, cluster density parameter, and contact risk parameter, and fusing the cluster size parameter, cluster density parameter, and contact risk parameter based on the first importance parameter to obtain the disease infection seat information.
[0102] In this embodiment, the cluster size parameter can be determined by the ratio of the number of occurrences of the target audio and historical target audio in the seat cluster area to the number of occurrences of the target audio and historical target audio in the target classroom. The cluster size parameter can be used to characterize the event scale characteristics of the cluster area. For example, in this embodiment, the cluster size parameter can specifically be the ratio of the total number of cough events in the cluster to the total number of cough events in the classroom within the time window. The numerical range of the cluster size parameter can be normalized to [0, 10].
[0103] In this embodiment, the cluster density parameter can be determined by the ratio of the number of occurrences of the target audio and historical target audio in the seat cluster area to the area of the seat cluster area. The cluster density parameter can be used to characterize the event density of the cluster area. For example, in this embodiment, the cluster density parameter can specifically be the ratio of the total number of cough events in the cluster to the coverage area of the cluster. The numerical range of the cluster density parameter can be normalized to [0, 10].
[0104] In this embodiment of the application, the contact risk parameter can be determined by the ratio of the number of at-risk students in the seating cluster area to the total number of students in the area. The contact risk parameter can be used to characterize the risk of infection from contact among people in the cluster area. For example, in this embodiment of the application, the contact risk parameter can specifically be the ratio of the number of students involved in the cluster to the total number of students in the cluster's coverage grid. The numerical range of the contact risk parameter can be normalized to [0, 10].
[0105] In the embodiments of this application, the first importance parameter can be a weighted average of the cluster size parameter, cluster density parameter, and contact risk parameter. This parameter can be used to determine the importance of each parameter in the fusion analysis. For example, the first importance parameter in the embodiments of this application can be specifically set according to the risk characteristics of respiratory infectious disease transmission in classrooms. The first importance parameter of the cluster size parameter can be set to 0.4, the first importance parameter of the cluster density parameter can be set to 0.3, and the first importance parameter of the contact risk parameter can be set to 0.3. Moreover, this weight can be flexibly adjusted according to the actual prevention and control needs of the campus and the characteristics of the classroom scenario.
[0106] In this embodiment of the application, the cluster size parameter, cluster density parameter, and contact risk parameter are fused based on the first importance parameter to obtain the disease infection seat information. The calculation formula of the classroom space clustering index (CSI) (i.e., the disease infection seat information in this embodiment of the application) is shown in Formula 16, where a can represent the weight of the cluster size parameter, Cs can represent the cluster size parameter, b can represent the weight of the cluster density parameter, Cd can represent the cluster density parameter, c can represent the weight of the contact risk parameter, and Cr can represent the contact risk parameter.
[0107] (Formula Sixteen) As an optional approach, when performing the "fusion analysis of time trend parameters, cumulative scale parameters, and type parameters corresponding to time feature information to obtain disease infection time information corresponding to at-risk students," the following methods can also be used, but are not limited to: determining time trend parameters based on the changing trend of time sequence in time feature information; determining cumulative scale parameters based on the frequency of occurrence of target audio and the number of students in the target classroom; determining the proportion of consecutive timestamps in the target timestamp as a type parameter; determining the second importance parameter of the time trend parameter, cumulative scale parameter, and type parameter, and fusing the time trend parameter, cumulative scale parameter, and type parameter based on the second importance parameter to obtain disease infection time information.
[0108] In this embodiment, the time trend parameter can be a quantitative parameter representing the trend of the infectious disease over time. The time trend parameter can be determined based on the time series changes of the target audio and historical target audio occurrences, and its numerical range can be normalized to [0, 10]. For example, if the increase in the number of occurrences of the target audio and historical target audio within three consecutive time windows is greater than or equal to 20%, the time trend can be an upward trend, and the time trend parameter can be 10; if the fluctuation range of the number of occurrences of the target audio and historical target audio within three consecutive time windows is less than or equal to 10%, the time trend can be a stable trend, and the time trend parameter can be 5; if the decrease in the number of occurrences of the target audio and historical target audio within three consecutive time windows is greater than or equal to 20%, the time trend can be a downward trend, and the time trend parameter can be 0.
[0109] In this embodiment of the application, the cumulative scale parameter can be the ratio of the number of times the target audio appears within a unit time window to the number of students in the target classroom. The cumulative scale parameter can be used to characterize the cumulative scale of disease infection over time, and the numerical range of the cumulative scale parameter can be normalized to [0, 10]. For example, in this embodiment of the application, the cumulative scale parameter can specifically be the ratio of the cumulative cough frequency in the current window to the total number of students in the classroom.
[0110] In the embodiments of this application, the type parameter can be the proportion of consecutive timestamps in the target timestamp. The type parameter can be used to represent the severity of symptoms of disease infection, and the numerical range of the type parameter can be normalized to [0, 10]. For example, in the embodiments of this application, the type parameter can specifically be the ratio of the number of consecutive cough events to the total number of cough events. Consecutive cough can be defined as the occurrence of more than or equal to 3 coughs within 10 seconds.
[0111] In this embodiment, the second importance parameter can be a weighted average of the time trend parameter, cumulative scale parameter, and type parameter. This second importance parameter can be used to determine the importance of each parameter in the temporal feature fusion analysis. For example, the second importance parameter in this embodiment can be set according to the temporal transmission patterns of respiratory infectious diseases. The second importance parameter for the time trend parameter can be 0.5, the second importance parameter for the cumulative scale parameter can be 0.3, and the second importance parameter for the type parameter can be 0.2. This weight can be dynamically adjusted based on the characteristics of prevalent infectious diseases on campus.
[0112] In this embodiment of the application, the time trend parameter, cumulative scale parameter and type parameter are fused based on the second importance parameter to obtain the disease infection time information. The calculation formula of the Time Transmission Index (TSI) (that is, the disease infection time information in this embodiment of the application) can be as shown in Formula 17, where d can represent the weight of the time trend parameter, Tt can represent the time trend parameter, e can represent the weight of the cumulative scale parameter, Ts can represent the cumulative scale parameter, f can represent the weight of the type parameter, and Tc can represent the type parameter.
[0113] (Formula 17) As an optional approach, when implementing the "disease transmission warning in the target classroom based on the disease infection seat information and disease infection time information in the disease infection information," the following methods can also be used, but are not limited to: fusing the disease infection seat information, disease infection time information, and air circulation parameters of the target classroom to obtain disease transmission risk data for the target classroom; fusing the number of students at risk and historically at risk, the number of areas with clustered seats, and disease transmission risk data to obtain comprehensive risk data; performing risk level analysis on the comprehensive risk data to obtain the disease transmission risk level corresponding to the comprehensive risk data; generating warning information corresponding to the disease transmission risk level based on the disease transmission risk level, and issuing a disease transmission warning in the target classroom based on the warning information. The warning information includes the second seat information of at-risk students, the target timestamp of the target audio, the disease infection seat information, the disease infection time information, the disease transmission risk data, the disease transmission risk level, and disease transmission prevention recommendations.
[0114] In this embodiment, the air circulation parameter can be a coefficient characterizing the ventilation conditions of the target classroom. The air circulation parameter can reflect the degree of influence of the classroom environment on the droplet transmission of respiratory infectious diseases. For example, in a classroom with both natural and air conditioning ventilation, the air circulation parameter can be 0.8; in a classroom with only natural ventilation, the air circulation parameter can be 0.5; and in a classroom without ventilation, the air circulation parameter can be 0.2. A smaller air circulation parameter indicates a longer droplet transmission time and a higher risk of respiratory infectious disease transmission.
[0115] In the embodiments of this application, disease transmission risk data can be data representing the risk of disease transmission. For example, the disease transmission risk data in the embodiments of this application can specifically be the Temporal-Spatial Communication Index (TSCI).
[0116] In this embodiment of the application, during the fusion processing, the maximum value of the classroom spatial clustering index (CSI) in the disease infection seat information and the current time transmission index (TSI) in the disease infection time information can be extracted. Then, combined with the influence factor (1-K) of the air circulation parameter, the disease transmission risk data of the target classroom can be calculated by weighted summation. The value range of the disease transmission risk data can be normalized to [0,10]. The larger the value of the disease transmission risk data, the higher the overall transmission risk.
[0117] In this embodiment of the application, the calculation formula for the Spatiotemporal Transmission Comprehensive Index (TSCI) (i.e., the disease transmission risk data in this embodiment of the application) can be as shown in Formula 18, where g can represent the weight of the maximum value of the classroom space aggregation index, CSImax can represent the maximum value of the classroom space aggregation index, h can represent the weight of the current time transmission index, TSIcurrent can represent the current time transmission index, i can represent the weight of the influence factor of the air circulation parameter, 1-K can represent the influence factor of the air circulation parameter, and K can represent the air circulation parameter.
[0118] (Formula 18) In this embodiment, the comprehensive risk data can be a quantitative indicator that integrates individual risk, regional risk, and overall transmission risk. This comprehensive risk data can be used to comprehensively assess the disease transmission risk within the target classroom. For example, the comprehensive risk data in this embodiment can specifically be the Virus Propagation Index (VPI).
[0119] For the embodiments of this application, the formula for calculating the Virus Transmission Index (VPI) can be as shown in Formula Nineteen, where j can represent the individual risk index weight, Vind can represent the individual risk index, n can represent the regional risk index weight, Varea can represent the regional risk index, l can represent the overall risk index weight, and Vtotal can represent the overall risk index; Vind can be determined based on the number of students at risk and students with historical risk, and the maximum frequency of the target audio emitted by students at risk. The formula for calculating Vind can be as shown in Formula Twenty, where A can represent the number of students at risk and students with historical risk, B can represent the total number of students in the target classroom, C can represent the maximum frequency of the target audio emitted by students at risk, and the value range of Vind can be [0, 100].
[0120] (Formula 19) (Formula 20) In this embodiment of the application, Varea can be determined based on the number of seat cluster areas and disease transmission risk data. The formula for calculating Varea can be as shown in Formula 21, where D can represent the number of seat cluster areas with high CSI, E can represent the total number of seat cluster areas, TSCImax can represent the maximum TSCI value, and the value range of Varea can be [0, 100].
[0121] (Formula 21) In this embodiment of the application, Vtotal can be determined based on disease transmission risk data and trend weights. The calculation formula of Vtotal can be as shown in Formula 22, where TSCI can represent the overall disease transmission risk data of the classroom, F can represent the trend weight, the trend weight of low-risk trends can be 0, the trend weight of medium-risk trends can be 0.5, the trend weight of high-risk trends can be 1, and the value range of Vtotal can be [0,100].
[0122] (Formula 22) In this embodiment of the application, the disease transmission risk level can be a risk level divided according to the threshold range of comprehensive risk data, and the disease transmission risk level can be used to trigger corresponding early warning operations. For example, the disease transmission risk level in this embodiment of the application can be divided into three levels. When the comprehensive risk data VPI is less than 30, the disease transmission risk level is Level 1 (low risk). Level 1 early warning indicates that there is only sporadic target audio in the classroom, without spatiotemporal clustering characteristics of infectious diseases, and the transmission risk is low. When the comprehensive risk data VPI is greater than or equal to 30 and less than 60, the disease transmission risk level is Level 2 (medium risk). Level 2 early warning indicates that there are spatiotemporal clustering characteristics of target audio in the classroom, indicating a risk of local infectious disease transmission, requiring targeted intervention. When the comprehensive risk data VPI is greater than or equal to 60, the disease transmission risk level is Level 3 (high risk). Level 3 early warning indicates that there are obvious spatiotemporal transmission characteristics of infectious diseases in the classroom, with a high transmission risk, requiring immediate and comprehensive prevention and control measures.
[0123] For example, each risk level can also have different characteristic conditions. A Level 1 warning can have no high CSI clustered areas, a TSI of less than 4, and the number of students at risk of emitting high-frequency target audio is less than or equal to 1. A Level 2 warning can have 12 high CSI clustered areas, a TSI greater than or equal to but less than 7, and the number of students at risk of emitting high-frequency target audio is 23, with no cross-regional spread. A Level 3 warning can have more than or equal to 3 high CSI clustered areas or 1 super high-rise clustered area, a TSI greater than or equal to 7, and the number of students at risk of emitting high-frequency target audio is more than or equal to 4, with cross-regional spread.
[0124] In the embodiments of this application, the recommendations for preventing the spread of disease can be targeted prevention and control suggestions based on the level of disease transmission risk and the spatiotemporal characteristics of its spread. For example, when the disease transmission risk level is a Level 1 warning, the recommendations for preventing the spread of disease could include daily ventilation and environmental disinfection; when the disease transmission risk level is a Level 2 warning, the recommendations could include targeted disinfection of clustered areas and adjusting the seating arrangement of students in adjacent seats; and when the disease transmission risk level is a Level 3 warning, the recommendations could include suspending classes, conducting health checks on affected students, and thoroughly disinfecting classrooms.
[0125] In this embodiment, after generating early warning information, tiered and regional early warning triggering operations can be executed according to the risk level. For example, when the disease transmission risk level is Level 1, the early warning information can be recorded only in the background; when the disease transmission risk level is Level 2, pop-up windows and SMS reminders can be sent to the homeroom teacher and school doctor; when the disease transmission risk level is Level 3, pop-up windows, SMS, and voice reminders can be sent to relevant personnel in charge of campus epidemic prevention and control; at the same time, the visualization platform can output panoramic heat maps of classrooms, risk indicator dashboards, spatiotemporal analysis reports, etc., to realize visualized disease transmission early warning, and after the early warning, it can enter dynamic monitoring mode to update risk data in real time and automatically adjust the early warning level.
[0126] As an optional approach, when performing the task of "acquiring audio and video data of the target classroom, performing voiceprint recognition on the target audio in the audio and video data, and identifying the students at risk of disease infection corresponding to the target audio", the following methods can also be used, but are not limited to: filtering audio data within the target frequency range from the audio data and identifying the selected audio data as the target audio; determining the target voiceprint information corresponding to the target audio and matching the target voiceprint information with the voiceprint information of multiple students to obtain voiceprint matching degree data; and identifying the students corresponding to the target voiceprint information as students at risk of disease infection if the voiceprint matching degree data meets the voiceprint matching conditions.
[0127] In this embodiment, the target frequency range can be a characteristic frequency band of vocalizations related to respiratory infectious diseases. For example, the target frequency range in this embodiment can specifically be 500-2000Hz, which is the core characteristic frequency band of vocalizations produced by the human nasal cavity and / or throat.
[0128] In the embodiments of this application, the audio data can be preprocessed before screening. The invalid audio signals such as table and chair movement and environmental noise are removed by using a combination of spectral subtraction and wavelet transform for noise reduction. At the same time, the audio data is framed and pre-emphasized.
[0129] In this embodiment, the target voiceprint information can be an acoustic feature vector extracted from the target audio. The target voiceprint information can be used to distinguish the vocal characteristics of different students. For example, the target voiceprint information in this embodiment can specifically include zero-crossing rate, energy, autocorrelation function in the time domain, and 12th-order Mel frequency cepstral coefficients + 1st-order energy, spectral center, and spectral bandwidth in the frequency domain. After extraction, Z-score normalization can be used to process the feature vector.
[0130] In this embodiment of the application, the voiceprint matching condition can be a voiceprint similarity threshold. For example, the voiceprint matching condition in this embodiment can be set to a cosine similarity greater than or equal to 90%. During voiceprint matching, a cosine similarity algorithm can be used, and the calculated cosine similarity value is the voiceprint matching score data. If the voiceprint matching score data meets the voiceprint similarity threshold, the match is considered successful, and the corresponding student is identified as a high-risk student. If the voiceprint matching score data does not meet the voiceprint similarity threshold, the target audio is determined to be invalid, and no further high-risk student identification operation is performed.
[0131] As an optional approach, this application also provides a multimodal map example. The multimodal map includes students, voiceprints, faces, and planar coordinates. The multimodal map can be used to integrate multi-dimensional information within a classroom, accurately locate and trace the risk of disease infection, and improve the accuracy of early warnings. The multimodal map can be based on the actual spatial layout of the target classroom, associating and mapping information such as students, voiceprints, faces, and planar coordinates. Specifically, student information can be associated with student identification, voiceprint information can be associated with the baseline voiceprint features of each student and the real-time voiceprint data corresponding to the target audio, face information can be associated with the student's face coordinates and facial movement features, and planar coordinates can be associated with the global physical coordinates of the classroom seats and the installation coordinates of the microphone and camera. Through the association of the above modal information, the correspondence between student identity, voiceprint features, face features, and seat positions can be achieved.
[0132] As an optional approach, this application embodiment also provides an example of tracing disease transmission paths. Transmission path tracing can be based on the spatiotemporal distribution characteristics of at-risk students, the time sequence of infection events, and seat associations to infer the disease transmission trajectory. Specific steps may include: extracting the second seat information of at-risk students, seat clustering areas from disease infection seat information, and time characteristic data from disease infection time information from a multimodal map to clarify the chronological order of infection time, seat adjacency, and clustering associations of each at-risk student; combining the timestamp of the target audio and the historical timestamp of historical target audio to trace the at-risk student whose target audio first appeared; based on the diffusion pattern of seat clustering areas and changes in time trend parameters, deriving the path of diffusion from the initial risk source to surrounding seats, including contact transmission paths between directly adjacent seats and group transmission paths within clustering areas; combining air circulation parameters to analyze the impact of ventilation conditions on the transmission path and determine the possible range and direction of droplet transmission; integrating the traced transmission path, initial risk source, and diffusion nodes (such as high-risk clustered seats) into the multimodal map and early warning information to further improve the scientific nature and efficiency of disease prevention and control.
[0133] Compared with existing technologies, this embodiment improves the accuracy of seat information consistency verification by determining the horizontal orientation and radial distance information of the target audio for sound source localization and refining seat information consistency verification; it accurately acquires student seat information by determining student facial coordinate information based on camera lens and image parameters and mapping it to obtain student seat information; it enhances the targeting of infection risk action recognition by matching target videos based on timestamps and identifying infection risk actions in candidate video frame sets; it achieves targeted early warning of disease transmission by analyzing seat clustering relationships and temporal feature information and fusing them to obtain spatiotemporal information of disease transmission; it improves the scientific nature of disease transmission early warning by fusing multi-dimensional parameters to obtain comprehensive risk data and generating graded early warning information; and it improves the accuracy of identifying students at risk of disease transmission by screening audio within the target frequency range and performing voiceprint matching.
[0134] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides an early warning device for the risk of disease infection, such as... Figure 4 As shown, the device includes: an acquisition module 31, a verification module 32, an identification module 33, and an early warning module 34.
[0135] The acquisition module 31 is configured to acquire audio and video data of the target classroom, perform voiceprint recognition on the target audio in the audio and video data, and identify the student at risk of disease infection corresponding to the target audio. The target audio is the sound made through the nasal cavity and / or throat when infected with a disease. The verification module 32 is configured to locate the sound source of the target audio based on the student seating information of the target classroom, obtain the first seat information corresponding to the target audio, and perform a consistency verification of the seat information based on the first seat information and the second seat information corresponding to the risk student. The identification module 33 is configured to identify the infection risk actions of students at risk based on the target video corresponding to the target audio in the audio and video data, when the first seat information and the second seat information meet the consistency condition. The infection risk actions include facial movements and / or limb movements. The early warning module 34 is configured to respond to the infection risk action of identifying at-risk students, determine the disease infection information corresponding to the at-risk students, and issue a disease transmission warning in the target classroom based on the disease infection information.
[0136] In some examples of this embodiment, the verification module 32 is specifically configured to determine the horizontal azimuth information and radial distance information between the sound source location of the target audio and the microphone in the target classroom, and to perform sound source localization on the target audio based on the horizontal azimuth information and radial distance information to obtain the sound source coordinate information of the target audio; based on the positioning accuracy of the sound source coordinate information, to determine the sound source range corresponding to the sound source coordinate information, and to select the first seat within the sound source range from multiple student seats corresponding to the target classroom based on the student seat information, and to determine the first seat information based on the first seat; to determine the second seat information corresponding to the at-risk student from the student seat information, and to perform a consistency verification of the seat information based on the first seat information and the second seat information to obtain the consistency verification result corresponding to the first seat information and the second seat information.
[0137] In some examples of this embodiment, the verification module 32 is further configured to determine the pixel offset angle of the camera relative to the student's face based on the lens parameters, image parameters, and center pixel coordinates of the camera in the target classroom; determine the student's face coordinate information in the target classroom based on the pixel offset angle, the height of the student's face from the ground, the camera's pose parameters, and the camera's global coordinates; and map the face coordinate information to the seating layout in the target classroom to obtain the student's seating information.
[0138] In some examples of this embodiment, the identification module 33 is further configured to, when the first seat information and the second seat information meet the consistency condition, determine the target timestamp of the target audio from the set of timestamps corresponding to the audio and video data; select video data corresponding to the target timestamp from the audio and video data, and determine the selected video data as the target video corresponding to the target audio, and determine a set of candidate video frames containing the action scenes of students at risk from the target video; and, based on the set of candidate video frames, perform facial action feature recognition on the facial actions of students at risk, and / or perform limb action feature recognition on the limb actions of students at risk, so as to identify whether there are infection risk actions in the set of candidate video frames that meet the conditions for disease infection actions.
[0139] In some examples of this embodiment, the early warning module 34 is specifically configured to perform seat adjacency analysis and seat clustering analysis based on the second seat information and the historical seat information of historically at-risk students to obtain the seat clustering area corresponding to the at-risk students; perform fusion analysis based on the clustering scale parameter, clustering density parameter, and contact risk parameter corresponding to the seat clustering area to obtain the disease infection seat information corresponding to the at-risk students; perform time sequence analysis and time interval analysis on the target timestamp and the historical target timestamp of the historical target audio to obtain the time feature information corresponding to the at-risk students; perform fusion analysis on the time trend parameter, cumulative scale parameter, and type parameter corresponding to the time feature information to obtain the disease infection time information corresponding to the at-risk students; and conduct disease transmission early warning in the target classroom based on the disease infection seat information and disease infection time information in the disease infection information.
[0140] In some examples of this embodiment, the early warning module 34 is further configured to: determine the cluster size parameter of the seat cluster area based on the number of occurrences of the target audio and historical target audio in the seat cluster area and the number of occurrences of the target audio and historical target audio in the target classroom; determine the cluster density parameter based on the number of occurrences of the target audio and historical target audio in the seat cluster area and the area of the seat cluster area; determine the contact risk parameter based on the number of at-risk students in the seat cluster area and the number of at-risk students in the target classroom; determine the first importance parameter of the cluster size parameter, the cluster density parameter and the contact risk parameter, and perform fusion processing on the cluster size parameter, the cluster density parameter and the contact risk parameter based on the first importance parameter to obtain disease infection seat information.
[0141] In some examples of this embodiment, the early warning module 34 is further configured to determine a time trend parameter based on the changing trend of the time sequence in the time feature information; determine a cumulative scale parameter based on the number of times the target audio appears and the number of students in the target classroom; determine the proportion of consecutive timestamps in the target timestamp as a type parameter; determine a second importance parameter for the time trend parameter, cumulative scale parameter and type parameter, and perform fusion processing on the time trend parameter, cumulative scale parameter and type parameter based on the second importance parameter to obtain disease infection time information.
[0142] In some examples of this embodiment, the early warning module 34 is further configured to fuse the disease infection seat information, disease infection time information, and air circulation parameters of the target classroom to obtain disease transmission risk data of the target classroom; fuse the number of students at risk and historically at risk, the number of areas where seats are clustered, and disease transmission risk data to obtain comprehensive risk data; perform risk level analysis on the comprehensive risk data to obtain the disease transmission risk level corresponding to the comprehensive risk data; generate early warning information corresponding to the disease transmission risk level based on the disease transmission risk level, and issue disease transmission warnings in the target classroom based on the early warning information. The early warning information includes the second seat information of the at-risk student, the target timestamp of the target audio, the disease infection seat information, the disease infection time information, the disease transmission risk data, the disease transmission risk level, and disease transmission prevention suggestions.
[0143] In some examples of this embodiment, the acquisition module 31 is specifically configured to filter out audio data within the target frequency range from the audio data and determine the selected audio data as the target audio; determine the target voiceprint information corresponding to the target audio and match the target voiceprint information with the voiceprint information of multiple students to obtain voiceprint matching degree data; if the voiceprint matching degree data meets the voiceprint matching conditions, determine the student corresponding to the target voiceprint information as a student at risk of disease infection.
[0144] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method shown.
[0145] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0146] like Figure 5 The diagram shown is a hardware structure schematic of an electronic device according to the present invention, comprising: At least one processor 401; and, Memory 402 is communicatively connected to at least one processor 401; wherein, The memory 402 stores instructions that can be executed by at least one processor to enable the at least one processor to perform the aforementioned early warning method for disease infection risk.
[0147] Figure 5 Take a processor 401 as an example.
[0148] The electronic device may also include an input device 403 and a display device 404.
[0149] The processor 401, memory 402, input device 403, and display device 404 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.
[0150] Memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the disease infection risk early warning method in the embodiments of this application, for example, Figure 1 and Figure 2 The method flow is shown. The processor 401 executes various functional applications and data processing by running non-volatile software programs, instructions, and modules stored in the memory 402, thereby realizing the disease infection risk early warning method in the above embodiments.
[0151] Memory 402 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the disease infection risk warning method, etc. Furthermore, memory 402 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 402 may optionally include memory remotely located relative to processor 401, and these remote memories may be connected via a network to the apparatus performing the disease infection risk warning method. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0152] Input device 403 can receive user clicks and signal inputs related to user settings and function controls for generating early warning methods for disease infection risks. Display device 404 may include display devices such as a display screen.
[0153] One or more modules are stored in memory 402, and when run by one or more processors 401, they execute the disease infection risk warning method in any of the above method embodiments.
[0154] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0155] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0156] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the solution of this embodiment, compared with the existing technology, this embodiment realizes multimodal infection risk verification of audio and video by identifying infection risk actions of at-risk students based on target video when the seat information meets the consistency conditions, thereby improving the accuracy of disease infection identification; by determining disease infection information after identifying infection risk actions and issuing disease transmission warnings in the target classroom, the effectiveness of classroom disease infection risk warnings is improved; by determining the horizontal orientation and radial distance information of the target audio for sound source localization and refining seat information consistency verification, the accuracy of seat information consistency verification is improved; by determining the student's risk based on camera lens, image and other parameters, the accuracy of seat information consistency verification is improved. The system maps facial coordinates to student seating information, enabling precise acquisition of student seating data. It improves the targeting of infection risk action identification by matching target videos based on timestamps and identifying infection risk actions from candidate video frames. It achieves targeted early warning of disease transmission by analyzing seating clustering relationships and temporal features and fusing them. It enhances the scientific rigor of disease transmission early warning by fusing multi-dimensional parameters to obtain comprehensive risk data and generating graded early warning information. Finally, it improves the accuracy of identifying students at risk of disease infection by filtering audio within a target frequency range and performing voiceprint matching.
[0158] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0159] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for early warning of disease infection risk, characterized in that, include: Acquire audio and video data during classroom teaching in the target classroom, perform voiceprint recognition on the target audio in the audio and video data, extract the voiceprint features of the target audio and match them with the voiceprint information of students in the target classroom, identify the students who emit the target audio, and determine the students at risk of disease infection corresponding to the target audio. The target audio is the sound emitted through the nasal cavity and / or throat when infected with a disease. Based on the student seating information of the target classroom, the target audio is localized to obtain the first seating information corresponding to the target audio. Then, based on the first seating information and the second seating information actually corresponding to the risk student in the target classroom, the consistency of the seating information is checked to determine whether the seat corresponding to the second seating information is within the sound source range corresponding to the first seating information. If the first seat information and the second seat information meet the consistency condition, the infection risk action identification is performed on the student at risk based on the target video corresponding to the target audio in the audio and video data. The infection risk action includes facial movements and / or limb movements. The step of identifying infection risk actions of the at-risk student based on the target video corresponding to the target audio in the audio and video data, when the first seat information and the second seat information meet the consistency condition, includes: determining the target timestamp of the target audio from the timestamp set corresponding to the audio and video data; selecting video data corresponding to the target timestamp from the audio and video data, and determining the selected video data as the target video corresponding to the target audio, and determining a candidate video frame set containing the action scenes of the at-risk student from the target video; and performing facial action feature recognition on the facial actions of the at-risk student and / or performing limb action feature recognition on the limb actions of the at-risk student based on the candidate video frame set, so as to identify whether there are infection risk actions in the candidate video frame set that meet the conditions for disease infection actions. In response to the identification of an infected student at risk, the system determines the disease infection information corresponding to the infected student and issues a disease transmission warning in the target classroom based on the disease infection information. This includes: analyzing seat adjacency and clustering relationships based on the second seat information and the historical seat information of historically at-risk students; analyzing the seat adjacency between the infected student and historically at-risk students and the spatial clustering of multiple infected students to obtain the seat clustering area corresponding to the infected student (the historically at-risk students are those previously identified as having a disease infection risk); performing a fusion analysis based on the clustering scale parameters, clustering density parameters, and contact risk parameters corresponding to the seat clustering area to obtain the disease infection seat information corresponding to the infected student, which is used to assess the spatial transmission risk of the disease; and performing a fusion analysis on the target timestamp and historical data. The historical timestamps of the target audio are analyzed for both temporal sequence and time interval to obtain the temporal characteristic information corresponding to the at-risk students. The temporal sequence analysis is the process of analyzing the chronological order of the target audio and the historical target audio, and is used to extract the temporal evolution pattern of disease infection. The time interval analysis is the process of analyzing the time interval between the target audio and the historical audio. The temporal characteristic information is used to quantify the temporal transmission characteristics of disease infection. The temporal trend parameters, cumulative scale parameters, and type parameters corresponding to the temporal characteristic information are fused and analyzed to obtain the disease infection time information corresponding to the at-risk students. The disease infection time information is used to assess the temporal transmission risk of the disease. Based on the disease infection seat information and the disease infection time information in the disease infection information, a disease transmission warning is issued in the target classroom. The method of conducting disease transmission early warning in the target classroom based on the disease infection seat information and the disease infection time information in the disease infection information includes: fusing the disease infection seat information, the disease infection time information, and the air circulation parameters of the target classroom to obtain disease transmission risk data of the target classroom. The air circulation parameters are coefficients characterizing the ventilation conditions of the target classroom and are used to reflect the degree of influence of the classroom environment on the droplet transmission of respiratory infectious diseases; and fusing the number of students at risk and students at historical risk, the number of areas where seats are clustered, and the disease transmission risk data to obtain comprehensive risk data. The comprehensive risk data is a quantitative indicator that integrates individual risk, regional risk, and overall transmission risk.
2. The method according to claim 1, characterized in that, The step of locating the sound source of the target audio based on the student seating information of the target classroom to obtain the first seating information corresponding to the target audio, and performing a consistency check on the seating information based on the first seating information and the second seating information actually corresponding to the at-risk student in the target classroom to determine whether the seat corresponding to the second seating information is within the sound source range corresponding to the first seating information, includes: Determine the horizontal azimuth and radial distance information between the sound source location of the target audio and the microphone in the target classroom, and perform sound source localization on the target audio based on the horizontal azimuth and radial distance information to obtain the sound source coordinate information of the target audio; Based on the positioning accuracy of the sound source coordinate information, the sound source range corresponding to the sound source coordinate information is determined, and the first seat within the sound source range is selected from multiple student seats corresponding to the target classroom according to the student seat information, and the first seat information is determined based on the first seat. The second seat information corresponding to the student at risk is determined from the student seat information, and the consistency of the seat information is checked based on the first seat information and the second seat information to obtain the consistency check result corresponding to the first seat information and the second seat information.
3. The method according to claim 2, characterized in that, Before determining whether the seat corresponding to the target audio is within the sound source range corresponding to the first seat information based on the student seating information of the target classroom, and before performing a consistency check on the seat information based on the first seat information and the second seat information actually corresponding to the at-risk student in the target classroom to determine whether the seat corresponding to the second seat information is within the sound source range corresponding to the first seat information, the method further includes: Based on the lens parameters, image parameters, and center pixel coordinates of the camera in the target classroom, the pixel offset angle of the camera relative to the student's face is determined; Based on the pixel offset angle, the height of the student's face from the ground, the camera's pose parameters, and the camera's global coordinates, the student's face coordinates in the target classroom are determined. The face coordinate information is mapped to the seating layout in the target classroom to obtain the student seating information.
4. The method according to claim 1, characterized in that, The method involves fusing and analyzing the cluster size parameters, cluster density parameters, and contact risk parameters corresponding to the seat clustering area to obtain the disease infection seat information corresponding to the at-risk students, including: Based on the frequency of occurrence of the target audio and historical target audio in the seat clustering area and the frequency of occurrence of the target audio and historical target audio in the target classroom, the clustering scale parameter of the seat clustering area is determined; The clustering density parameter is determined based on the frequency of occurrence of the target audio and historical target audio in the seat clustering area and the area of the seat clustering area. The contact risk parameters are determined based on the number of at-risk students in the seating area and the number of at-risk students in the target classroom. The first importance parameter of the cluster size parameter, the cluster density parameter, and the contact risk parameter is determined, and the cluster size parameter, the cluster density parameter, and the contact risk parameter are fused based on the first importance parameter to obtain the disease infection seat information.
5. The method according to claim 1, characterized in that, The step of fusing and analyzing the time trend parameters, cumulative scale parameters, and type parameters corresponding to the time feature information to obtain the disease infection time information corresponding to the at-risk students includes: The time trend parameter is determined based on the changing trend of the time sequence in the time feature information; The cumulative scale parameter is determined based on the number of times the target audio appears and the number of students in the target classroom; The proportion of consecutive timestamps in the target timestamp is determined as the type parameter; A second importance parameter is determined for the time trend parameter, the cumulative scale parameter, and the type parameter, and the time trend parameter, the cumulative scale parameter, and the type parameter are fused based on the second importance parameter to obtain the disease infection time information.
6. The method according to claim 1, characterized in that, The step of issuing a disease transmission warning in the target classroom based on the disease infection seat information and the disease infection time information in the disease infection information includes: Risk level analysis is performed on the comprehensive risk data to obtain the disease transmission risk level corresponding to the comprehensive risk data; Based on the disease transmission risk level, a warning message corresponding to the disease transmission risk level is generated, and a disease transmission warning is issued in the target classroom according to the warning message. The warning message includes the second seat information of the at-risk student, the target timestamp of the target audio, the disease infection seat information, the disease infection time information, the disease transmission risk data, the disease transmission risk level, and disease transmission prevention suggestions.
7. The method according to claim 1, characterized in that, The process of acquiring audio and video data from classroom teaching in the target classroom, performing voiceprint recognition on the target audio in the audio and video data, extracting the voiceprint features of the target audio and matching them with the voiceprint information of students in the target classroom, identifying the student who emitted the target audio, and determining the student at risk of disease infection corresponding to the target audio, includes: Audio data within the target frequency range is selected from the audio and video data, and the selected audio data is determined as the target audio. Determine the target voiceprint information corresponding to the target audio, and match the target voiceprint information with the voiceprint information of multiple students to obtain voiceprint matching degree data; If the voiceprint matching data meets the voiceprint matching conditions, the student corresponding to the target voiceprint information is identified as a student at risk of disease infection.
8. A disease infection risk early warning device, characterized in that, include: The acquisition module is configured to acquire audio and video data during classroom teaching in the target classroom, perform voiceprint recognition on the target audio in the audio and video data, extract the voiceprint features of the target audio and match them with the voiceprint information of students in the target classroom, identify the students who emit the target audio, and determine the students at risk of disease infection corresponding to the target audio, wherein the target audio is a sound emitted through the nasal cavity and / or throat in the case of disease infection. The verification module is configured to locate the sound source of the target audio based on the student seating information of the target classroom, obtain the first seating information corresponding to the target audio, and perform a consistency verification of the seating information based on the first seating information and the second seating information actually corresponding to the risk student in the target classroom, and determine whether the seat corresponding to the second seating information is within the sound source range corresponding to the first seating information. The identification module is configured to identify infection risk actions of the at-risk student based on the target video corresponding to the target audio in the audio and video data, when the first seat information and the second seat information meet the consistency condition. The infection risk actions include facial movements and / or limb movements. The identification module is further configured to, when the first seat information and the second seat information meet the consistency condition, identify the infection risk actions of the at-risk student based on the target video corresponding to the target audio in the audio and video data, including: when the first seat information and the second seat information meet the consistency condition, determining the target timestamp of the target audio from the timestamp set corresponding to the audio and video data; selecting video data corresponding to the target timestamp from the audio and video data, and determining the selected video data as the target video corresponding to the target audio, and determining a candidate video frame set containing the action scenes of the at-risk student from the target video; and, based on the candidate video frame set, performing facial action feature recognition on the at-risk student's facial actions, and / or performing limb action feature recognition on the at-risk student's limb actions, to identify whether there are infection risk actions in the candidate video frame set that meet the conditions for disease infection actions; The early warning module is configured to respond to the identification of an infection risk action of a student at risk, determine the disease infection information corresponding to the student at risk, and issue a disease transmission warning in the target classroom based on the disease infection information. This includes: performing seat adjacency analysis and seat clustering analysis based on the second seat information and the historical seat information of historically at-risk students; analyzing the seat adjacency between the student at risk and the historically at-risk students, and the spatial clustering of multiple students at risk, to obtain the seat clustering area corresponding to the student at risk, where the historically at-risk students are those previously identified as having a disease infection risk; performing a fusion analysis based on the clustering scale parameters, clustering density parameters, and contact risk parameters corresponding to the seat clustering area to obtain the disease infection seat information corresponding to the student at risk, which is used to assess the spatial transmission risk of the disease; and issuing a disease transmission warning in the target classroom. The time sequence analysis and time interval analysis are performed on the time stamps and historical target timestamps of the target audio to obtain the time feature information corresponding to the at-risk students. The time sequence analysis is the process of analyzing the chronological order of the target audio and the historical target audio, and is used to extract the temporal evolution pattern of disease infection. The time interval analysis is the process of analyzing the time interval between the target audio and the historical audio. The time feature information is used to quantify the temporal transmission characteristics of disease infection. The time trend parameters, cumulative scale parameters, and type parameters corresponding to the time feature information are fused and analyzed to obtain the disease infection time information corresponding to the at-risk students. The disease infection time information is used to assess the temporal transmission risk of the disease. Based on the disease infection seat information and the disease infection time information in the disease infection information, a disease transmission warning is issued in the target classroom. The early warning module is further configured to issue a disease transmission early warning in the target classroom based on the disease infection seat information and the disease infection time information in the disease infection information. This includes: fusing the disease infection seat information, the disease infection time information, and the air circulation parameters of the target classroom to obtain disease transmission risk data for the target classroom. The air circulation parameters are coefficients characterizing the ventilation conditions of the target classroom and are used to reflect the degree of influence of the classroom environment on the droplet transmission of respiratory infectious diseases; and fusing the number of students at risk and students with historical risk, the number of areas where seats are clustered, and the disease transmission risk data to obtain comprehensive risk data. The comprehensive risk data is a quantitative indicator that integrates individual risk, regional risk, and overall transmission risk.
Citation Information
Patent Citations
Infectious disease early warning method and device based on artificial intelligence, electronic equipment and medium
CN114724730A
Identity recognition system, method and device
CN115221488A