Audio and video acquisition module, and audio and video acquisition device
By using multi-camera layout design and image processing technology, the problem of student positioning and behavior analysis in the recording and broadcasting solution was solved, realizing accurate positioning and behavior analysis of students in the classroom, supporting 3D modeling and face recognition, and improving the accuracy of classroom data analysis.
Patent Information
- Application Number
- PCT/CN2024/101036
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-24
- Publication Date
- 2026-01-02
AI Technical Summary
Existing audio and video recording solutions cannot perform classroom scene analysis, teaching data production, and classroom feedback based on the collected audio and video data. They cannot achieve accurate positioning and behavior analysis of students in the classroom, especially when the camera's view is obstructed during student activities, making it impossible to accurately identify student positions.
The system employs a multi-camera layout design, including at least one first camera device and at least two second camera devices, which respectively collect dynamic information and spatial location sub-information of the object. The spatial location information of the object is calculated by the image processing unit and matched with audio data for analysis, thereby achieving accurate positioning and behavioral analysis of students.
It enables precise positioning and behavioral analysis of students in the classroom, accurately identifies student locations during student activities, supports 3D modeling and facial recognition, and improves the accuracy of classroom data analysis and teaching effectiveness.
Smart Images

Figure CN2024101036_02012026_PF_FP_ABST
Abstract
Description
Audio and video acquisition module and audio and video acquisition device TECHNICAL FIELD
[0001] The present application relates to the technical field of electronic products, in particular to an audio and video acquisition module and an audio and video acquisition device. BACKGROUND
[0002] With the development of smart home, smart teaching equipment and smart conference equipment, network camera devices and network microphone devices have gradually become popular. For example, in the informationization deployment of school classrooms, network camera devices and network microphone devices can record the classroom content such as the classroom behavior of teachers and students, the blackboard writing of teachers and the classroom performance of teacher-student interaction in the form of audio and video. That is, network camera devices and network microphone devices can be used by schools in specific courses, such as public courses, remote courses, etc. In related technologies, the implementation of such similar scene functions mainly adopts a recording and broadcasting scheme, but the current recording and broadcasting audio and video scheme still has many problems. It only completes the real recording of the classroom to achieve the effect of restoring and presenting the audio and video, and it cannot provide more in-depth and effective audio and video data for classroom scene analysis, teaching data production, classroom feedback, etc. based on the collected audio and video.
[0003] SUMMARY
[0004] Therefore, it is necessary to provide an audio and video acquisition module and an audio and video acquisition device for acquiring audio and video data, so that the audio and video data collected by the audio and video acquisition module can be used to accurately analyze the object activity information in the target scene.
[0005] In a first aspect, the present application provides an audio and video acquisition module deployed in a target space, comprising:
[0006] A first image acquisition unit, the first image acquisition unit comprises at least one first camera device, each first camera device is used for acquiring object dynamic information in a respective field of view angle;
[0007] A second image acquisition unit, the second image acquisition unit comprises at least two second camera devices, each second camera device is used for acquiring object spatial position sub-information in a respective field of view angle; wherein the union set of the field of view angle of the first camera device and the union set of the field of view angle of the second camera device match the target space;
[0008] a first image processing unit and a second image processing unit, an input end of the first image processing unit being electrically connected with an output end of the first image acquisition unit, and an input end of the second image processing unit being electrically connected with an output end of the second image acquisition unit; the second image processing unit is configured to receive and process the object spatial position sub-information of the object; the second image processing unit determines the object spatial position information of the object in the target space based on at least two object spatial position sub-information of the same object, position information of each second image capturing device in the target space corresponding to each object spatial position sub-information, and angle information of an imaging plane of each corresponding second image capturing device; the first image processing unit is configured to receive and process the object dynamic information of the object; and the behavior analysis of the object is realized by the object spatial position information and the object dynamic information.
[0009] In one of the embodiments, the system further comprises:
[0010] an audio acquisition unit configured to acquire audio data in the target space;
[0011] an audio processing unit, an input end of the audio processing unit being electrically connected with an output end of the audio acquisition unit, configured to receive and process the audio data; and an output end of the audio processing unit being electrically connected with the first image processing unit and the second image processing unit respectively, configured to transmit the processed audio data to the first image processing unit and / or the second image processing unit, so as to realize data matching of the processed audio data in the target space with the processed object dynamic information and / or the object spatial position information.
[0012] In one of the embodiments, the system further comprises:
[0013] a data switch, input ends of the data switch being electrically connected with output ends of the first image processing unit, the second image processing unit and the audio processing unit respectively, configured to receive encoded data; wherein the encoded data comprises the processed object dynamic information, the object spatial position information and / or the audio data;
[0014] a communication interface, electrically connected with an output end of the data switch, configured to realize external transmission of the encoded data.
[0015] In one of the embodiments, the communication interface is multiplexed as a power supply interface.
[0016] In one of the embodiments, the input end of the audio processing unit is also electrically connected with the output end of another audio collecting unit arranged in the target space; wherein the another audio collecting unit is also used to collect audio data in the target space.
[0017] In one of the embodiments, at least two of the audio collecting devices in the audio collecting unit are connected in cascade.
[0018] In one of the embodiments, the field of view angle of the first camera device is determined based on at least the width of the target space and the maximum distance between the target space and the corresponding first camera device.
[0019] The field of view angle of the second camera device is determined based on the width of the target space and the maximum distance between the target space and the corresponding second camera device.
[0020] In one of the embodiments, at least two of the second camera devices in the second image collecting unit are arranged on both sides of at least one of the first camera devices in the same direction.
[0021] In one of the embodiments, the first camera device comprises a long-focus camera device.
[0022] The second camera device comprises a wide-angle camera device.
[0023] In one of the embodiments, the first image collecting unit comprises at least three first camera devices, and the three first camera devices are symmetrically arranged.
[0024] In one of the embodiments, the field of view angle of each two of the first camera devices partially overlaps.
[0025] In one of the embodiments, the second image collecting unit comprises at least one set of binocular camera devices.
[0026] In a second aspect, the present application further provides an audio-video collecting device comprising the audio-video collecting module.
[0027] The audio-video collecting device further comprises a display screen.
[0028] The audio and video acquisition module and the audio and video acquisition device provided by the present application are characterized in that the audio and video acquisition module is arranged in a target space and includes at least a first image acquisition unit, a second image acquisition unit, a first image processing unit and a second image processing unit. The first image acquisition unit includes at least one first camera device, and each first camera device is used to acquire dynamic information of an object in a respective field of view. The second image acquisition unit includes at least two second camera devices, and each second camera device is used to acquire spatial position sub-information of an object in a respective field of view. The union of the fields of view of the first camera devices and the union of the fields of view of the second camera devices match the target space. The input end of the first image processing unit is electrically connected to the output end of the first image acquisition unit, and the input end of the second image processing unit is electrically connected to the output end of the second image acquisition unit. The second image processing unit is used to receive and process the acquired spatial position sub-information of an object. The second image processing unit determines the spatial position information of an object in the target space based on at least two pieces of spatial position sub-information of the same object, the position information of each second camera device in the target space corresponding to each piece of spatial position sub-information, and the angle information of the imaging surface of each corresponding second camera device. The first image processing unit is used to receive and process the acquired dynamic information of an object. The behavior of an object is analyzed based on the spatial position information and the dynamic information.
[0029] As can be seen, this application sets up an audio and video acquisition module with image acquisition units of different specifications to acquire object dynamic information and object spatial location information in the target space respectively, and sets up different image processing units to process object dynamic information and object spatial location information respectively. This allows different types of information to be acquired by camera devices of different specifications and processed by different image processing units, which helps to improve the acquisition and processing effects of different types of information. This makes the acquired different types of information more accurate and clear, and the data obtained after processing by the corresponding image processing units is more accurate and can more accurately reflect the behavior, performance and interaction information of objects in the target space. Furthermore, this application sets the number of second camera devices in the audio and video acquisition module for acquiring object spatial location information within the target space to at least two, so that the second image processing unit can acquire at least two object spatial location sub-information for the same object within the target space. By calculating the spatial coordinate difference corresponding to the object spatial location sub-information acquired for the target object, as well as the specific installation position and installation angle of each second camera device, the accurate object spatial location information of the target object in the target space can be obtained. That is, the audio and video acquisition module provided by this application can improve the accuracy of the finally determined object spatial location information of the object within the target space based on at least two object spatial location sub-information corresponding to the same object, as well as the position information and imaging plane angle information of each second camera device corresponding to each object spatial location sub-information within the target space. In addition, this application uses at least two second camera devices to acquire the object spatial location sub-information of the same target object, realizing that the finally determined object spatial location information of the target object is three-dimensional information, further improving the accuracy of the determined object spatial location information. Furthermore, since this case targets the same object, at least two second camera devices are used to collect its spatial location sub-information. Based on these at least two object spatial location sub-information, the object spatial location information of the target object is determined. Therefore, even if other objects or obstacles obscure the target object during the process of capturing its spatial location sub-information, as long as the second camera devices can collect some of the target object's position coordinates, the specific object spatial location information of the target object can be calculated using the calculation method provided in this application. This allows for the identification of the target object's accurate location in the target space, which is beneficial for subsequent accurate analysis of the target object's behavior information within the target space based on the original video data. Attached Figure Description
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only represent some of the embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without any creative effort.
[0031] Fig. 1 shows a structural schematic diagram of an audio and video acquisition module provided by an embodiment of the present application;
[0032] Fig. 2 shows a principle diagram of two second camera devices shooting the same object provided by an embodiment of the present application;
[0033] Fig. 3 shows a target space diagram including other audio acquisition units provided by an embodiment of the present application;
[0034] Fig. 4 shows a calculation principle diagram of a field of view angle provided by an embodiment of the present application;
[0035] Fig. 5 shows a layout schematic diagram of an image acquisition unit provided by an embodiment of the present application;
[0036] Fig. 6 shows a field of view angle schematic diagram of a first image acquisition unit provided by an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to facilitate the understanding of the present application, the following will make a more comprehensive description of the present application with reference to the relevant drawings. The drawings show the preferred embodiments of the present application. However, the present application can be realized in many different forms, and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive.
[0038] It should be noted that when an element is referred to as being "fixed" to another element, it can be directly on the other element or there can be an intervening element. When an element is referred to as being "connected" to another element, it can be directly connected to the other element or intervening elements can be present. As used herein the term "vertical", "horizontal", "left", "right", and the like are merely used for the purpose of illustration.
[0039] In this document, spatially relative terms, such as "upper" and "lower", are defined with respect to the drawings. Thus, "upper" and "lower" can be used interchangeably with respect to one another. It will be understood that when a layer is referred to as being "on" another layer, it can be directly formed on the other layer or intervening layers can also be present. Therefore, when a layer is referred to as being "directly on" another layer, there are no intervening layers present between the two layers.
[0040] In the drawings, the size of layers and regions can be exaggerated for clarity. It will be understood that when a layer or element is referred to as being "on" another layer or substrate, it can be directly on the other layer or substrate, or intervening layers can also be present. Also, it will be understood that when a layer is referred to as being "between" two layers, it can be the only layer between the two layers, or one or more intervening layers can also be present. In addition, like reference numerals are used throughout the drawings to denote like elements.
[0041] Hereinafter, although terms such as "first", "second", etc. can be used to describe various components, the components are not necessarily limited to the above terms. The above terms are used only to distinguish one component from another component. It will also be understood that an expression used in the singular encompasses an expression in the plural, unless the context clearly dictates otherwise. In addition, in the following embodiments, it will also be understood that the terms "comprise" and / or "have" used herein indicate the presence of the stated features or components, but do not exclude the presence or addition of one or more other features or components.
[0042] In the following embodiments, when a layer, region, or element is "connected", it can be interpreted as not only being directly connected but also being connected through another constituent element placed therebetween. For example, when a layer, region, element, etc. is described as being connected or electrically connected, the layer, region, element, etc. can be connected or electrically connected not only directly but also through another layer, region, element, etc. placed therebetween.
[0043] As used in the specification, the term "and / or" includes any and all combinations of one or more of the associated listed items. When phrases such as "at least one of (a), (b), and (c)" are used, it is meant to include any one of (a), (b), or (c) individually, as well as any combination of (a), (b), and (c).
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the specification herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0045] It will also be understood that the terms "comprise / comprising" or "have / having" etc. specify the presence of stated features, integers, steps, operations, components, parts, or combinations thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, components, parts, or combinations thereof.
[0046] With the continuous development of school classrooms towards intelligence and technology, the requirements for recorded classroom audio and video are also increasing, for example: the need for face recognition of each student in the classroom, and the recording of student classroom behavior, etc. In the current school classroom informatization deployment, network camera devices and network microphone devices are gradually popularized. These devices can record the classroom behavior of teachers and students, the teacher's board writing, and the classroom performance of teacher-student interaction in the form of audio and video. Currently, the recording function of audio and video information similar to the classroom is mainly realized through some recording and broadcasting schemes. When obtaining audio and video data, dynamic information of students and spatial position information of the classroom need to be collected. However, the traditional recording and broadcasting scheme is limited by the fixed field of view design and resolution of the camera; for example, the camera has insufficient shooting accuracy, and can only realize simple audio and video recording, and cannot perform face recognition on students in the classroom and locate students; thus, it is impossible to accurately associate the classroom behavior (such as raising hands, interaction, etc.) of students with the students themselves based on the recorded video. Therefore, the traditional recording and broadcasting scheme cannot collect dynamic information of students and spatial position information of students.
[0047] In view of the above situation, the traditional classroom observation device upgrades the hardware, mainly by replacing the specifications of the camera or changing the arrangement of the existing camera. However, due to the influence of the classroom scene, students are concentrated in a specific area of the classroom, and when the camera collects image information, the students in the back row are often blocked by the students in the front row. In this case, although the camera has the level of collecting dynamic information of students and spatial position information of the classroom, the focus cannot be achieved due to the blocking of part of the view angle during collection, which ultimately leads to blurred video and cannot accurately identify the position of the student himself, so as to analyze the classroom behavior of the student himself.
[0048] Therefore, the present application first changes the number of cameras used to collect classroom spatial information, and re-arranges all cameras used to collect classroom spatial information. One or more cameras are used to shoot the target student at the same time, and the spatial coordinate difference collected by each camera for the target student and the installation position of each camera are calculated to obtain the position information of the target student. The above-mentioned method of the present application can obtain the coordinates of the target student by calculation as long as the camera can collect part of the position coordinates of the target student when the target student is blocked by a student or an obstacle, so as to identify the position of the student himself and analyze the classroom behavior of the student himself.
[0049] Based on the problems in the above related technologies, the present application provides an audio and video acquisition module and an audio and video acquisition device for acquiring audio and video data, so that the audio and video data acquired by the audio and video acquisition module can be used to realize accurate positioning and accurate analysis of object activity information in a target scene.
[0050] Fig. 1 shows a structural schematic diagram of an audio and video acquisition module provided by an embodiment of the present application. Referring to Fig. 1, the present application provides an audio and video acquisition module 100 deployed in a target space, which comprises:
[0051] A first image acquisition unit 11, the first image acquisition unit 11 comprises at least one first camera device, and each first camera device is used at least to acquire object dynamic information in a respective field of view angle;
[0052] A second image acquisition unit 12, the second image acquisition unit 12 comprises at least two second camera devices, and each second camera device is used at least to acquire object spatial position sub-information in a respective field of view angle; wherein the union set of the field of view angles of the first camera devices and the union set of the field of view angles of the second camera devices are matched to the target space;
[0053] A first image processing unit 21 and a second image processing unit 22, the input end of the first image processing unit 21 is electrically connected to the output end of the first image acquisition unit 11, and the input end of the second image processing unit 22 is electrically connected to the output end of the second image acquisition unit 12; the second image processing unit 22 is used to receive and process the acquired object spatial position sub-information; wherein the second image processing unit 22 determines object spatial position information of an object in the target space based on at least two object spatial position sub-information of the same object, position information of each second camera device in the target space corresponding to each object spatial position sub-information, and angle information of an imaging surface of the corresponding each second camera device; the first image processing unit 21 is used to receive and process object dynamic information of the acquired object; and behavior analysis of the object is realized through the object spatial position information and the object dynamic information.
[0054] It should be explained that the target space in which the audio and video acquisition module 100 is deployed can be, for example, an indoor space such as a classroom, a conference room, a theater, etc., but the present application is not limited thereto, and the audio and video acquisition module 100 can also be deployed in an outdoor space.
[0055] The object in the target space can be any person in the target space, such as a student and a teacher in a classroom, or a participant in a conference room. However, the present application is not limited thereto, and under the demand circumstances, the object in the target space can also include some objects and animals.
[0056] Specifically, the application provides an alternative implementation, which is provided with an audio and video collection module 100, which at least includes a first image collection unit 11, a second image collection unit 12, a first image processing unit 21 and a second image processing unit 22. The first image collection unit 11 can be provided with at least one first camera device, or a plurality of first camera devices. Each first camera device in the first image collection unit 11 can be used to collect dynamic information of an object in its own field of view, such as action information, hand-raising information, interactive information and the like of a person object. The second image collection unit 12 can be provided with at least two second camera devices, or a larger number of second camera devices. Each second camera device in the second image collection unit 12 can be used to collect spatial position sub-information of an object in its own field of view, such as three-dimensional position coordinate sub-information of the collected object in a target space, i.e. spatial position of the object and front-back relationship of the object in the target space.
[0057] The union of the fields of view of all the first camera devices included in the first image collection unit 11 matches the entire target space, i.e. through all the first camera devices included in the first image collection unit 11, picture information of every required camera shooting position in the target space can be collected, such as picture information of all positions in the target space, or picture information of more than a preset proportion (such as 95% or 85%) of positions in the target space, to ensure the requirement of the first image collection unit 11 for collecting dynamic information of an object in the target space. Similarly, the union of the fields of view of all the second camera devices included in the second image collection unit 12 also matches the entire target space, to ensure the requirement of the second image collection unit 12 for collecting spatial position information of an object in the target space.
[0058] The first camera device and the second camera device can both be a camera, but in the technical solution provided by the application, the types of the object information collected by the first camera device and the second camera device are different, so the first camera and the second camera can be selectively provided with different specifications to respectively meet the requirement of the first camera for collecting dynamic information of an object in the target space and the requirement of the second camera for collecting spatial position sub-information of an object in the target space. The differentiated specifications are beneficial to improving the collection effect of different new types of information, and are also beneficial to controlling the cost of different specifications of camera devices.
[0059] Further, an input end of the first image processing unit 21 is electrically connected with an output end of the first image acquisition unit 11 (not shown). Here, the output end of each first camera device can be electrically connected with the input end of the first image processing unit 21 respectively, or the output end of the whole first image acquisition unit 11 can be electrically connected with the input end of the first image processing unit 21, which is not limited in the present application. The first image processing unit 21 is for example an image processing chip, which is used to receive and process the object dynamic information of the object collected by each first camera device, so as to at least realize the encoding of the picture data corresponding to the object dynamic information, and other processing operations can be performed on the received object dynamic information based on requirements.
[0060] An input end of the second image processing unit 22 is electrically connected with an output end of the second image acquisition unit 12 (not shown). Here, the output end of each second camera device can be electrically connected with the input end of the second image processing unit 22 respectively, or the output end of the whole second image acquisition unit 12 can be electrically connected with the input end of the second image processing unit 22, which is not limited in the present application. The second image processing unit 22 is for example an image processing chip, which is used to receive and process the object spatial position sub-information of the object collected by each second camera device, so as to at least realize the encoding of the picture data corresponding to the object spatial position sub-information, and other processing operations can be performed on the received object spatial position sub-information based on requirements. Further, the second image processing unit 22 provided in the present application can determine the specific object spatial position information of the object in the target space based on the at least two object spatial position sub-information of the same object, the position information of each second camera device corresponding to each object spatial position sub-information of the object in the target space, and the angle information of the imaging plane of each corresponding second camera device. That is, the specific position information of an object in the target space (object spatial position information) is not determined based on the picture data collected by one second camera device, but is determined based on the picture data of at least two second camera devices corresponding to the same object, combined with the position information of the at least two second camera devices in the target space and the installation angle information.
[0061] This application uses two or more second camera devices to capture images of the front-to-back positional relationships of human objects within a target space scene. After processing by the second image processing unit 22, the images can identify the 3D depth information, i.e., the three-dimensional coordinate information, of the human objects within the target space. The specific principle can be seen in Figure 2, which is a schematic diagram illustrating the principle of two second camera devices capturing the same object according to an embodiment of this application. When two cameras (second camera devices) are used to capture the same object (P) at a certain distance and angle, the coordinates of the object generated by the two cameras (left imaging plane and right imaging plane) are different. Based on these two coordinates and the positions and angles of the two cameras, the distance from the object to the two cameras can be calculated, thereby obtaining the specific three-dimensional coordinate position information of the object within the target space. It should be noted that this application does not specifically limit the conversion formulas or models used for the relevant conversions, as long as the effect of determining the object's spatial position information within the target space based on the collected object spatial position sub-information of the same object, the corresponding position information of the second camera device within the target space, and the angle information of the imaging plane is achieved.
[0062] The application sets different specifications of image acquisition units in the audio and video acquisition module 100 to respectively acquire object dynamic information and object space position sub-information in the target space, and sets different image processing units to respectively process the object dynamic information and the object space position sub-information, so that different types of information can be collected by different specifications of camera equipment, and processed by different image processing units, which is beneficial to improve the collection effect and processing effect of different types of information, so that the collected different types of information are more accurate and clear, and the data obtained after processing by the corresponding image processing unit is more accurate, and the behavior, performance and interaction information of the object in the target space can be more accurately fed back. Furthermore, the different specifications of image acquisition units correspond to different image processing units, which can improve the processing efficiency of image data compared with only using one image processing unit. In addition, the number of the second camera equipment in the audio and video acquisition module 100 for collecting object space position sub-information in the target space is at least two, so that the second image processing unit 22 can collect at least two object space position sub-information of the same object in the target space, and obtain the accurate object space position information of the target object in the target space by calculating the space coordinate difference of the object space position sub-information collected for the target object, and the specific installation position and installation angle of each second camera equipment; that is, the audio and video acquisition module provided by the application can improve the accuracy of the object space position information of the object in the target space based on at least two object space position sub-information of the same object, and the position information and angle information of the imaging surface of each second camera equipment in the target space; in addition, at least two second camera equipment are used to collect the object space position sub-information of the same target object, so that the object space position information of the target object is a three-dimensional coordinate position information, which further improves the accuracy of the determined object space position information; so as to realize the 3D modeling of multiple objects in the target space included in the video information based on the object space position information of each object, and combine the high-definition object face information collected by the first camera equipment to satisfy the matching of the specific object in the target space and the corresponding space position information.
[0063] It should be further noted that, since the present case is directed to the same target object, the object space position sub-information of the target object is collected by at least two second camera devices, and then the object space position information of the target object is determined based on the at least two object space position sub-information. Therefore, even in the process of collecting the object space position sub-information of the target object, as long as part of the position coordinates of the target object can be collected by the second camera device when other objects or obstacles block the target object, the specific object space position information of the target object can be calculated by the calculation method provided in the present application, so that the accurate position of the target object in the target space can be recognized, and then the behavior information of the target object in the target space based on the original video data can be accurately analyzed.
[0064] Please continue to refer to FIG. 1. In an exemplary embodiment, the audio and video collection module 100 further comprises:
[0065] The audio collection unit 13 is configured to collect audio data in the target space.
[0066] The audio processing unit 23 is electrically connected to the output end of the audio collection unit 13, configured to receive and process the audio data. The output end of the audio processing unit 23 is electrically connected to the first image processing unit 21 and the second image processing unit 22, configured to transmit the processed audio data to the first image processing unit 21 and / or the second image processing unit 22, so as to realize data matching between the processed audio data in the target space and the processed object dynamic information and / or object space position information.
[0067] Specifically, in the audio and video collection module 100 provided in the present application, in addition to the above-mentioned image collection units (11 and 12) and image processing units (21 and 22), the audio and video collection module 100 can further comprise an audio collection unit 13 and an audio processing unit 23. The input end of the audio processing unit 23 can be electrically connected to the output end of the audio collection unit 13. The present application does not make specific limitation on the number of audio collection devices included in the audio collection unit 13. In the case that two or more audio collection devices are arranged in the audio collection unit 13, the output end of each audio collection device can be electrically connected to the audio processing unit 23, or the output end of the entire audio collection unit 13 and the input end of the audio processing unit 23 can be electrically connected. The audio processing unit 23 is configured to receive and process the audio data collected by the audio collection unit 13 in the target space. The audio processing unit 23 can be an audio processing chip, for example, which is configured to at least encode the received audio data, so as to facilitate further transmission and storage of the audio data.
[0068] In the case that the first image processing unit 21, the second image processing unit 22 and the audio processing unit 23 are simultaneously arranged in the audio-video acquisition module 100, the application further provides an alternative embodiment that the output end of the audio processing unit 23 is electrically connected with the first image processing unit 21 and the second image processing unit 22 respectively (not shown), so that the audio data processed by the audio processing unit 23 can be transmitted to the first image processing unit 21 and / or the second image processing unit 22, to realize the matching of the audio data processed by the audio processing unit 23 with the object dynamic information and / or the object spatial position information processed by the image processing unit (21 and 22) in the target space, so as to realize the matching of the video picture and the audio data.
[0069] It should be noted that the above-mentioned arrangement that the output end of the audio processing unit 23 is electrically connected with the first image processing unit 21 and the second image processing unit 22 respectively is only one alternative embodiment provided by the application, and alternatively, as shown in FIG. 1, the output end of the audio processing unit 23 is electrically connected with the first image processing unit 21, and the second image processing unit 22 is further electrically connected with the first image processing unit 21, so that the audio data processed by the audio processing unit 23 can be further transmitted to the second image processing unit 22 by the first image processing unit 21.
[0070] The audio and video acquisition module 100 provided by the application can be used to record all-round classroom audio and video original data when applied in a teaching space such as a classroom. The audio data in the classroom range is recorded by the audio acquisition unit 13. The face picture in the classroom range is recorded by the first image acquisition unit 11. The 3D position information picture of the students and teachers in the classroom range is recorded by the second image acquisition unit 12, so that the collected audio and video data can realize 3D modeling and meet the accuracy requirement of face recognition. The first image processing unit 21, the second image processing unit 22 and the audio processing unit 23 are combined to accurately collect and process the position information of the students before and after the classroom, and further realize processing such as teaching board content extraction, face recognition and depth position information conversion. Then, the audio and video data processed by the audio and video acquisition module 100 provided by the application can be transmitted to the outside through a network signal, to provide the basic data and part of the processed data associated with the objects in the classroom to the outside, for example, the part of the data can be further given to the application back end for further function development, to realize correct recognition of the spatial position of each student in the classroom, and the face recognition algorithm corresponds to each student, so that the students in each area in the classroom, the state and interaction of each student in the teaching activity can be accurately presented and analyzed, so that the teacher can accurately receive feedback information such as the student's classroom enthusiasm according to the related analysis report, so that the teacher can effectively use the data to adjust the teaching scheme. In addition, the function developed by the back end may, for example, be used to help the school evaluate the classroom performance of each student and the teaching level of the teacher with objective data, thereby facilitating the improvement of the teacher's individualized teaching level for each student and helping the teacher improve the teaching ability.
[0071] The application also provides an alternative embodiment, as shown in FIG. 1, the audio and video acquisition module 100 includes three first camera devices, two second camera devices and eight audio acquisition devices (microphones), but this is only an alternative embodiment provided by the application, the application is not limited to this, and the application does not limit the arrangement order and arrangement position of the devices. The arrangement order and arrangement position of the devices can be selectively adjusted according to the requirements.
[0072] Please continue to refer to FIG. 1, in an exemplary embodiment, further comprising:
[0073] The data switch 24 is electrically connected to the output ends of the first image processing unit 21, the second image processing unit 22 and the audio processing unit 23 respectively, and is used to receive encoded data; wherein the encoded data includes processed object dynamic information, object spatial position information and / or audio data;
[0074] The communication interface 25 is electrically connected with the output end of the data switch 24, and is used for realizing the sending of the encoded data.
[0075] Specifically, the application further provides a selectable setting mode of the audio and video acquisition module 100, which is provided with the first image acquisition unit 11, the second image acquisition unit 12, the audio acquisition unit 13, the first image processing unit 21, the second image processing unit 22 and the audio processing unit 23, and is further provided with the data switch 24 and the communication interface 25, wherein the input end of the data switch 24 is electrically connected with the output end of the first image processing unit 21, the second image processing unit 22 and the audio processing unit 23 respectively, so as to receive the data processed by the three data processing units, and the data can be the encoded data processed by the three data processing units, i.e. the object dynamic information processed and encoded by the first image processing unit 21, the object spatial position information processed and encoded by the second image processing unit 22 and the audio data processed and encoded by the audio processing unit 23.
[0076] The above-mentioned setting mode that the input end of the data switch 24 is electrically connected with the output end of the first image processing unit 21, the second image processing unit 22 and the audio processing unit 23 respectively is only a selectable embodiment provided by the application, and if the original audio data and picture data do not need to be stored in actual use, the audio processing unit 23 and the first image processing unit 21 can be electrically connected, the first image processing unit 21 and the second image processing unit 22 can be electrically connected, and the second image processing unit 22 can be further electrically connected with the data switch 24, so that all the picture data and audio data processed by the second image processing unit 22 can be transmitted to the data switch 24, i.e. only the second image processing unit 22 and the data switch 24 can be electrically connected.
[0077] The communication interface 25 can be electrically connected with the output end of the data switch 24, so as to receive the encoded data received by the data switch 24 and realize the sending of the encoded data to an external terminal, which can be an external computer or an external server, so as to realize the processing of the picture data and the audio data included in the encoded data by the external computer or the external server, and realize the accurate analysis or accurate simulation of the information related to the object in the target space based on the picture data and the audio data acquired by the audio and video acquisition module 100.
[0078] Please continue to refer to Fig. 1, in an exemplary embodiment, the communication interface 25 is multiplexed as a power supply interface.
[0079] Specifically, the audio-video collection module 100 can further include a power interface (not shown) for realizing access of external power to realize normal work of the audio-video collection module 100. The power interface can be directly electrically connected with the data switch 24, and the power is transmitted to other unit modules in the audio-video collection module 100 based on the data switch 24. Here, an alternative embodiment is provided, that is, the communication interface 25 can be multiplexed as the power interface, so that the number of interfaces required to be arranged in the audio-video collection module 100 can be reduced, thereby facilitating simplification of the module architecture and reduction of the manufacturing cost of the module.
[0080] The audio-video collection module 100 provided in the application has audio data collection and processing, video data collection and processing, power supply and network transmission according to functions. The audio-video collection module 100 can be optionally applied in a school classroom. In order to realize that the audio-video data can cover the whole classroom, the frame can be selected and designed according to the size of the classroom, the pixel requirement of face recognition, the 3D depth design requirement, the installation and deployment convenience and other factors, so that clear sound pickup, full dead angle-free shooting picture and image data meeting the requirements of face recognition and 3D algorithm can be achieved.
[0081] The power interface of the application supports POE (Power Over Ethernet) and 12V power supply, and the data transmission part can be transmitted through the network. The POE can include a power sourcing equipment (PSE) and a power device (PD). The audio data and the video data can be transmitted through different network segments, and network segment isolation is beneficial to guarantee the stability of the data bandwidth. The communication interface 25 is multiplexed as the power interface, so that one network cable can be used for power supply and data transmission.
[0082] Fig. 3 shows a target space provided by an embodiment of the application and including other audio collection units. Please refer to Fig. 3 in combination with Fig. 1. In an exemplary embodiment, the input end of the audio processing unit 23 is further electrically connected with the output end of other audio collection units 31 arranged in the target space. The other audio collection units 31 are also used for collecting audio data in the target space.
[0083] Specifically, in addition to the structural elements included in the audio-video acquisition module 100 provided by the present application, other audio acquisition units 13 can be further arranged in the target space. For example, the audio-video acquisition module 100 provided by the present application is installed on one side wall in the target space, and the other audio acquisition units 13 are installed on the other side walls in the target space. Based on this, the present application further provides an alternative embodiment that the input end of the audio processing unit 23 in the audio-video acquisition module 100 is also electrically connected to the output end of the other audio acquisition units 31 arranged in the target space, so as to receive the audio data in the target space collected by the other audio acquisition units 31 at the same time when receiving the audio data in the target space collected by the audio acquisition unit 13 in the audio-video acquisition module 100. Whether it is the audio acquisition unit 13 in the audio-video acquisition module 100 or the other audio acquisition units 31 in the target space, the audio data collected by them is the audio data in the target space. Therefore, the audio data collected by the two types of devices can be transmitted to the audio processing unit 23 in the audio-video acquisition module 100, so that the clarity and accuracy of the audio data received by the audio processing unit 23 are higher, thereby facilitating the improvement of the audio acquisition effect of the audio-video acquisition module 100 and the accuracy of the subsequent analysis of the information related to the objects in the target space through the audio data.
[0084] Please refer to FIG. 1 and FIG. 3. In an exemplary embodiment, at least two audio acquisition devices in the audio acquisition unit 13 are connected in cascade.
[0085] Specifically, for the plurality of audio acquisition devices included in the audio acquisition unit 13, the present application provides an alternative embodiment that the plurality of audio acquisition devices included in the audio acquisition unit 13 are connected in cascade. Further, the plurality of audio acquisition devices included in the audio acquisition unit 13 and the plurality of other audio acquisition devices in the other audio acquisition units 31 in the target space can also be connected in cascade.
[0086] The audio acquisition device is, for example, a microphone. By connecting a plurality of microphones to an audio source in cascade, all the microphones connected in cascade can receive the same audio signal, which is beneficial to expand the voice pickup range and enhance the quality of the picked-up audio.
[0087] Further, the audio acquisition device arranged in the target space can be a smart microphone. Specifically, the smart microphone can pick up the sound according to the position of the sound object in the target space. For example, in the case that three smart microphones are arranged in sequence along the horizontal direction, when there is a sound object on the left side, the sound emitted by the sound object can be picked up by the smart microphone on the left side or the smart microphones on the left side and the middle side, which is beneficial to improve the sound pickup accuracy of the audio acquisition device in the audio acquisition unit 13 and improve the sound pickup effect. In addition, the smart microphone with sound tracking capability can be selected as the audio acquisition device, which is beneficial to further improve the sound pickup accuracy and sound pickup effect of the audio acquisition device. The smart microphone with sound tracking capability can move according to the position of the sound object, so that the sound receiving surface of the smart microphone faces the sound object.
[0088] In the case that the audio and video acquisition module 100 is applied to a classroom, in order to meet the sound pickup in the classroom, the audio acquisition unit 13 can be arranged with an 8-array microphone as shown in FIG. 1, and the electrical connection between the 8-array microphones can be realized by selecting a mode with cascading function, so as to realize the sound pickup of a larger angle and a longer distance in the classroom by the audio acquisition unit 13. The 8-array microphone can meet the sound pickup of more than 180°. In addition, the audio data of other sound pickup devices in other audio acquisition units 31 at the back end of the classroom can be processed by audio algorithm together, which can meet the clear sound pickup of the entire classroom.
[0089] When the audio and video acquisition module 100 is applied to a classroom environment, the face recognition can be realized, the personnel information of students and the classroom activity situation can be one-to-one corresponding, the 3D depth recognition can be realized, the student area can be accurately divided, the teaching feedback can be accurately to the group or even the individual area, the multiple module mode can be innovatively used in the teaching scene, the collected audio and video basic information can be used for correct feedback of classroom teaching in addition to the ordinary classroom record, and one network cable can be used for power supply and data transmission, which simplifies the deployment wiring.
[0090] Through the audio and video acquisition module 100, the collected audio data can be processed by a related algorithm to generate an audio file of the entire classroom, the video data can be first associated with the face recognition data to realize the correspondence between the picture and the person, and then the 3D data can be used to construct the spatial position of each person, so as to accurately restore the position of the student, make the information of each student in each position correspond, provide basic information for the subsequent application of the classroom interaction heat map of each position, the situation of each area and each student, and thus provide more accurate analysis data and suggestions for the teacher.
[0091] FIG. 4 is a diagram illustrating a principle of calculating a field of view according to an embodiment of the present application. Please refer to FIG. 1, FIG. 3 and FIG. 4. In an exemplary embodiment, the field of view of the first camera device is determined based on at least the width W of the target space and the maximum distance L between the corresponding first camera device and the target space.
[0092] The field of view of the second camera device is determined based on the width W of the target space and the maximum distance L between the corresponding second camera device and the target space.
[0093] Specifically, the camera devices shown in FIG. 4 include the first camera device and the second camera device. For the specification selection of the first camera device and the second camera device, the present application provides an alternative embodiment that the field of view of the corresponding camera device is determined based on at least the width of the target space where the audio and video collection module 100 is arranged and the maximum distance between the corresponding camera device and the target space.
[0094] In this embodiment, the audio and video collection module 100 is used in a classroom. The formula for calculating the horizontal field of view is tan(0.5θ) = 0.5W / L, where W is the width of the classroom and L is the maximum distance between the camera device and the farthest side wall in the classroom. The formula for calculating the vertical field of view is tan(0.5θ) = 0.5H / L, where H is the height of the classroom and L is the maximum distance between the camera device and the farthest side wall in the classroom. θ is the angle of the field of view.
[0095] For example, the width of the classroom ranges from 6 to 8 meters and the length ranges from 8 to 10 meters. To meet the requirement of shooting the picture in the classroom, two camera devices with the specifications of 8M, HFOV 130° and VFOV 80° can be selected for the second camera device. 8M represents 8 million pixels, HFOV represents the horizontal field of view, and VFOV represents the vertical field of view.
[0096] Since the first camera needs to collect the face information in the classroom, the pixel points of the face in the picture need to be greater than 64*64 to ensure the clarity of the face information, and the size of the student's face is generally 160*160mm; according to the principle of the camera, under the condition that the size of the shooting object is unchanged, the farther the object is from the camera, the smaller the pixel points it occupies; therefore, the requirement for the camera is that the pixel points of the 160mm*160mm face at a distance of 9m from the camera are greater than 64*64. Therefore, three first camera devices can be selected to set, and the technical solution of multiple camera splicing is adopted to realize the accuracy requirement of collecting face information, specifically, three 8M cameras are adopted, and the field of view angle of each camera is HFOV48°, VFOV43°, so that the pixel points of the face to be shot in the entire classroom are greater than 64*64, and the clarity of the face picture is ensured. For example, the field of view angle of the first camera device can be determined based on the width W of the target space and the maximum distance L between the corresponding second camera device in the target space and the width size D of the object. Here, the width size D of the object is, for example, the face width or the body width of the student.
[0097] Fig. 5 shows a schematic diagram of an image acquisition unit provided by an embodiment of the present application. Please refer to Fig. 1 and Fig. 5, in an exemplary embodiment, at least two second camera devices in the second image acquisition unit 12 are arranged on both sides of at least one first camera device in the same direction.
[0098] Specifically, since the second camera device is used for image acquisition of the same object, the same object needs to have different object space position sub-information, therefore, there needs to be a certain spacing space between the two second camera devices arranged adjacently; based on this, in the case that the audio and video acquisition module 100 includes three first camera devices and two second camera devices, the two second camera devices can be selected to be arranged on both sides of the three first camera devices; but the present application is not limited thereto, and the two second camera devices can also be selected to be arranged on both sides of the two first camera devices, or the two second camera devices can be selected to be arranged on both sides of the one first camera device, and if the spacing space is allowed, the first camera device can also be selected not to be arranged between the two second camera devices.
[0099] For the second image acquisition unit including two second camera devices, the two second camera devices are arranged at different positions, and when the two second camera devices shoot the object at the same position in the target space, the pictures shot by the two second camera devices have a certain parallax; wherein the parallax refers to the position difference of the same pixel (object) in the two pictures. The depth (distance) Z between the object and the second camera device is determined based on the parallax d, the spacing b between the two second camera devices and the focal length f of the second camera device, and the specific calculation formula is Z=(b*f) / d.
[0100] As shown in the implementation principle of FIG. 2, for two second camera devices, which can be regarded as a left camera and a right camera, the left camera is used to capture a target object, and a corresponding left imaging surface exists. The right camera is used to capture the same target object, and a corresponding right imaging surface exists. FIG. 2 takes the target object as point P for example, the left imaging point of point P on the left imaging surface is PL, and the right imaging point of point P on the right imaging surface is PR. The coordinates of point PL on the left imaging surface and the coordinates of point PR on the right imaging surface are different. Based on the coordinates of the target object on the left imaging surface and the right imaging surface, the setting position and the setting angle of the left and right cameras, the distance (depth) from the target object to the camera can be converted, and the position information of the target object in the target space is obtained. The actual calculation of the related position involves spatial coordinates and 3D algorithm, which is not limited and described in detail herein. In order to achieve better recognition accuracy, the distance between the two second camera devices can be set to 150 mm or 200 mm.
[0101] Please refer to FIG. 1. In an exemplary embodiment, the first camera device includes a long-focus camera device.
[0102] The second camera device includes a wide-angle camera device.
[0103] Specifically, for the types of the first camera device and the second camera device, an alternative embodiment provided by the present application is that the first camera device adopts a long-focus camera, and the second camera device adopts a wide-angle camera. The long-focus camera is used to realize high-resolution image acquisition with a large field of view, so as to realize face recognition based on image acquisition. The wide-angle camera is used to acquire object space position sub-information. Based on at least two object space position sub-information acquired by at least two wide-angle cameras, in combination with the position information of the wide-angle camera and the angle information of the imaging surface, the accurate object space position information (three-dimensional position coordinate information) of the object in the target space is determined.
[0104] Please refer to FIG. 1 and FIG. 5. In an exemplary embodiment, the first image acquisition unit 11 includes at least three first camera devices, and the three first camera devices are symmetrically arranged.
[0105] Among them, the field of view of each two first camera devices partially overlaps.
[0106] Specifically, in the case that the audio and video acquisition module 100 includes three first camera devices, the three first camera devices can be symmetrically arranged. For example, the three first camera devices are arranged in sequence along the same direction, and the middle first camera device is located on the symmetry axis.
[0107] Fig. 6 shows a schematic view of a field of view of a first image acquisition unit according to an embodiment of the present application. Please refer to Fig. 1, Fig. 5 and Fig. 6. In order to ensure that the first image acquisition unit 11 can acquire the object dynamic information of all objects in the target space, the field of view of every two first camera devices can be partially overlapped. In this way, the situation that some areas in the target space are not captured by any first camera device can be avoided. In an optional implementation of the present application, the field of view of the two first camera devices on the side can be arranged to have an optical axis angle with the lens surface of the middle camera device. For example, the optical axis angle can be 41.5 degrees. However, the present application is not limited in this regard.
[0108] Please refer to Fig. 1. In an exemplary embodiment, the second image acquisition unit 12 includes at least a group of binocular camera devices.
[0109] Specifically, in an optional implementation of the present application, the second image acquisition unit 12 can include at least a group of binocular camera devices in addition to the two second camera devices or more second camera devices. The two second camera devices with a certain distance can capture the same object position information, and the binocular camera devices can also capture the same object position information. That is, the binocular camera devices can capture the object position information of the same object in the target space.
[0110] Based on the same application concept, the present application further provides an audio and video acquisition device (not shown) including an audio and video acquisition module 100.
[0111] The audio and video acquisition device further includes a display screen.
[0112] Specifically, the audio and video acquisition device can be a smart blackboard. The audio and video acquisition module 100 provided by the present application can be integrated into the audio and video acquisition device. The smart blackboard includes a display screen for teachers' writing on the board and knowledge display.
[0113] In addition, the audio and video acquisition device can also be a smart conference machine, a teacher machine or a student machine, etc. The embodiments disclosed in the present application are not limited in this regard.
[0114] The audio and video acquisition module 100 integrated into the audio and video acquisition device is only an optional implementation of the present application. The audio and video acquisition module 100 can also be a separate audio and video acquisition device.
[0115] Any combination of the technical features in the above-described embodiments can be made, and for the sake of brevity, not all possible combinations are described, however, as long as there is no conflict, any combination of the technical features should be considered within the scope of the present disclosure.
[0116] The above-described embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the application. It should be pointed out that for ordinary skilled persons in the art, some modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. An audio / video acquisition module, characterized in that, Deployed within the target space, including: A first image acquisition unit, comprising at least one first camera device, wherein each first camera device is used to acquire dynamic information of an object within its respective field of view. The second image acquisition unit includes at least two second camera devices, each of which is used to acquire at least the spatial position sub-information of the object within its respective field of view; wherein the union of the field of view of the first camera device and the union of the field of view of the second camera device are both matched to the target space; A first image processing unit and a second image processing unit are connected, with the input terminal of the first image processing unit electrically connected to the output terminal of the first image acquisition unit, and the input terminal of the second image processing unit electrically connected to the output terminal of the second image acquisition unit. The second image processing unit is used to receive and process the acquired object spatial location sub-information of the object. The second image processing unit determines the object spatial location information of the object in the target space based on at least two object spatial location sub-information pieces for the same object, the position information of each second camera device corresponding to each object spatial location sub-information within the target space, and the angle information of the imaging plane of each corresponding second camera device. The first image processing unit is used to receive and process the acquired object dynamic information of the object, and to perform behavior analysis on the object through the object spatial location information and the object dynamic information.
2. The audio and video acquisition module according to claim 1, characterized in that, Also includes: An audio acquisition unit, wherein the audio acquisition unit is used to acquire audio data within the target space; An audio processing unit, wherein the input terminal of the audio processing unit is electrically connected to the output terminal of the audio acquisition unit, is used to receive and process the audio data; and the output terminal of the audio processing unit is electrically connected to the first image processing unit and the second image processing unit respectively, for transmitting the processed audio data to the first image processing unit and / or the second image processing unit, so as to achieve data matching between the processed audio data in the target space and the processed object dynamic information and / or the object spatial location information.
3. The audio and video acquisition module according to claim 2, characterized in that, Also includes: A data switch, wherein the input terminals of the data switch are electrically connected to the output terminals of the first image processing unit, the second image processing unit, and the audio processing unit, respectively, for receiving encoded data; wherein the encoded data includes processed object dynamic information, object spatial location information, and / or the audio data; The communication interface is electrically connected to the output of the data switch and is used to transmit the encoded data.
4. The audio and video acquisition module according to claim 3, characterized in that, The communication interface is reused as a power interface.
5. The audio and video acquisition module according to claim 2, characterized in that, The input terminal of the audio processing unit is also electrically connected to the output terminal of other audio acquisition units deployed in the target space; wherein, the other audio acquisition units are also used to acquire audio data in the target space.
6. The audio and video acquisition module according to claim 2, characterized in that, At least two audio acquisition devices in the audio acquisition unit are cascaded together.
7. The audio and video acquisition module according to claim 1, characterized in that, The field of view of the first camera device is determined at least based on the width of the target space and the maximum distance between the target space and the corresponding first camera device; The field of view of the second camera device is determined based on the width of the target space and the maximum distance between the target space and the corresponding second camera device.
8. The audio and video acquisition module according to claim 1, characterized in that, Along the same direction, at least two of the second camera devices in the second image acquisition unit are respectively located on both sides of at least one of the first camera devices.
9. The audio and video acquisition module according to claim 8, characterized in that, The first camera device includes a telephoto camera device; The second camera device includes a wide-angle camera device.
10. The audio and video acquisition module according to claim 8, characterized in that, The first image acquisition unit includes at least three first camera devices, which are arranged symmetrically. In this case, the field of view of each pair of the first camera devices partially overlaps.
11. The audio and video acquisition module according to claim 1, characterized in that, The second image acquisition unit includes at least one set of binocular camera devices.
12. An audio and video acquisition device, characterized in that, Includes the audio and video acquisition module as described in any one of claims 1-11; The audio and video acquisition device also includes a display screen.
Citation Information
Patent Citations
Video processing method and device
CN110505403A
Image detection method and device
CN113673503A
Shooting control method and video acquisition system
CN117097990A
Distributed video monitoring system
JP1997130783A