Intelligent picture-in-picture display method and device for multi-video stream analysis, equipment and medium
By identifying speakers and dividing the display area, the problem of unclear display of speakers in multi-camera meetings is solved, and a clear picture-in-picture display and overall visual effect are achieved, making it easier for speakers to adjust their speeches.
Patent Information
- Application Number
- CN202510552298.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-11
AI Technical Summary
In multi-camera meeting occasions, the prior art cannot effectively identify speakers and display them clearly on the screen, resulting in inconvenience to observe speakers and poor overall visual effects.
By obtaining deployment space information and camera location, identifying speakers and determining target cameras, dividing display on display screens, realizing picture-in-picture display, automatically identifying and displaying speakers and overall visual effects.
It is convenient for intuitive observation of speakers, helps speakers grasp the overall atmosphere and adjust the rhythm of speech, and provides a clear picture-in-picture display effect.
Smart Images

Figure CN120302146A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video, and in particular, to an intelligent picture-in-picture display method, apparatus, device, and medium for multi-video stream analysis. Background Art
[0002] With the development of camera technology, cameras are more and more widely used. For example, some conference occasions may involve the use of multiple cameras. For example, when there are multiple cameras in a conference room, usually one camera is used to shoot one person, and then the video streams captured by each camera are displayed on the screen. However, since there is usually only one speaker, when each captured video stream is displayed on the screen, it is not conducive for people other than the speaker to quickly find the speaker, and since the screen displays multiple video streams, the display ratio of the speaker on the screen is small, which is also not conducive to viewing. Summary of the Invention
[0003] Embodiments of this application provide an intelligent picture-in-picture display method, apparatus, device, and medium for multi-video stream analysis to solve at least one problem existing in the related art. The technical solutions are as follows:
[0004] In a first aspect, embodiments of this application provide a method for intelligent picture-in-picture display of multi-video stream analysis, including:
[0005] Obtain the spatial information of the deployment space, the position information of each camera in the deployment space, and the video streams of each camera;
[0006] Determine a first target camera according to each of the position information and the spatial information, where the first target camera is used to provide an overall visual effect;
[0007] Perform speaker recognition on each of the video streams to obtain a speaker recognition result, and determine a second target camera according to the speaker recognition result, where the second target camera is used to shoot the speaker;
[0008] Determine the main screen area and the picture-in-picture area of the display screen, display the video stream of the first target camera in the picture-in-picture area, and display the video stream of the second target camera in the main screen area.
[0009] In an implementation manner, the determining the first target camera according to each of the position information and the spatial information includes:
[0010] When the spatial information includes a central position, according to each piece of the position information and the central position, determine a first distance between each camera and the central position, and use the camera with the smallest first distance among the movable cameras as the first target camera, or use the camera with the largest first distance as the first target camera, or use the camera with the smallest first distance as the first target camera;
[0011] Or,
[0012] When the spatial information includes a specific position tag, according to each piece of the position information and the specific position tag, determine a second distance between each camera and the specific position tag, and use the camera with the smallest second distance as the first target camera.
[0013] In one implementation, the performing speaker recognition on each of the video streams to obtain a speaker recognition result includes:
[0014] Perform face recognition on each of the video streams to determine whether there is a face in each video stream;
[0015] Perform pose recognition on the video streams with faces to determine whether there is a speaker in a speaking pose, and obtain a speaker recognition result.
[0016] In one implementation, the determining a second target camera according to the speaker recognition result includes:
[0017] When the speaker recognition result is that there is a speaker in a speaking pose, determine the camera corresponding to the video stream in which the speaker is recognized as the second target camera;
[0018] When the speaker recognition result is that there is no speaker in a speaking pose, use the camera corresponding to the video stream with a default face as the second target camera, or determine the second target camera from the cameras corresponding to the video streams with faces.
[0019] In one implementation, the method further includes:
[0020] When a speaker is recognized in the video stream of the second target camera, track the ID continuity of the speaker through a tracking algorithm, where when a speaker is recognized, the ID of the speaker will be recorded;
[0021] When the ID continuity of the speaker is characterized as inconsistent, match the ID of the speaker with the IDs in the video streams of other cameras except the current second target camera, and update the camera corresponding to the video stream with a successful match as the new second target camera.
[0022] In one embodiment, the determining the main screen area and the picture-in-picture area of the display screen includes:
[0023] Obtain the screen resolution of the display screen to determine the main screen area, and determine whether there is a preset picture-in-picture ratio;
[0024] When there is a preset picture-in-picture ratio, determine the picture-in-picture area according to the product of the preset picture-in-picture ratio and the screen resolution;
[0025] When there is no preset picture-in-picture ratio, based on the default picture-in-picture ratio or the picture-in-picture ratio input after prompting, determine the picture-in-picture area according to the product of the picture-in-picture ratio and the screen resolution.
[0026] In one embodiment, the method further includes:
[0027] When the video stream of the first target camera is displayed in the picture-in-picture area and the video stream of the second target camera is displayed in the main screen area, it includes at least one of the following:
[0028] In response to a click instruction of the user on the picture-in-picture area, switch the picture-in-picture area and the main screen area;
[0029] In response to a drag instruction of the user on the picture-in-picture area, move the picture-in-picture area or adjust the size of the picture-in-picture area.
[0030] In a second aspect, an intelligent picture-in-picture display device for multi-video stream analysis provided by an embodiment of the present application includes:
[0031] An acquisition module, configured to acquire the spatial information of the deployment space, the position information of each camera in the deployment space, and the video stream of each camera;
[0032] A first determination module, configured to determine a first target camera according to each of the position information and the spatial information, where the first target camera is used to provide an overall visual effect;
[0033] A second determination module, configured to perform speaker recognition on each of the video streams to obtain a speaker recognition result, and determine a second target camera according to the speaker recognition result, where the second target camera is used to capture the speaker;
[0034] A display module, configured to determine the main screen area and the picture-in-picture area of the display screen, display the video stream of the first target camera in the picture-in-picture area, and display the video stream of the second target camera in the main screen area.
[0035] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory. Instructions are stored in the memory and loaded and executed by the processor to implement the method in any one of the above aspects.
[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it implements the method in any one of the above aspects.
[0037] The beneficial effects in the above technical solutions at least include:
[0038] By obtaining the spatial information of the deployment space, the position information of each camera in the deployment space, and the video streams of each camera, determining the first target camera according to each position information and the spatial information, performing speaker recognition on each video stream to obtain a speaker recognition result, and determining the second target camera according to the speaker recognition result, where the second target camera is used to capture the speaker, determining the main screen area and the picture-in-picture area of the display screen, displaying the video stream of the first target camera in the picture-in-picture area, and displaying the video stream of the second target camera in the main screen area, automatically recognizing the speaker and displaying it in the main screen area, which is convenient for intuitively and clearly observing the speaker. At the same time, displaying the video stream of the first target camera with the overall visual effect in the picture-in-picture area is beneficial for the speaker to grasp the overall atmosphere and adjust the rhythm and content of the speech.
[0039] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the above-described illustrative aspects, embodiments, and features, further aspects, embodiments, and features of the present application will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in the present application and should not be regarded as limiting the scope of the present application.
[0041] Figure 1 It is a schematic flowchart of the steps of a method for intelligent picture-in-picture display of multi-video stream analysis according to an embodiment of the present application;
[0042] Figure 2 It is a structural block diagram of an intelligent picture-in-picture display device for multi-video stream analysis according to an embodiment of the present application;
[0043] Figure 3 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] In the following, only some exemplary embodiments are simply described. As can be recognized by those skilled in the art, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the accompanying drawings and the description are considered to be exemplary in nature and not restrictive.
[0045] Referring to Figure 1 , a flowchart of an intelligent picture-in-picture display method for multi-video stream analysis according to an embodiment of the present application is shown. The intelligent picture-in-picture display method for multi-video stream analysis may at least include steps S100 - S400:
[0046] S100. Obtain the spatial information of the deployment space, the position information of each camera in the deployment space, and the video stream of each camera.
[0047] S200. Determine a first target camera according to each position information and the spatial information, where the first target camera is used to provide an overall visual effect.
[0048] S300. Perform speaker recognition on each video stream to obtain a speaker recognition result, and determine a second target camera according to the speaker recognition result, where the second target camera is used to capture the speaker.
[0049] S400. Determine the main screen area and the picture-in-picture area of the display screen, display the video stream of the first target camera in the picture-in-picture area, and display the video stream of the second target camera in the main screen area.
[0050] The intelligent picture-in-picture display method for multi-video stream analysis according to the embodiment of the present application can be executed by an electronic control unit, a controller, a processor, etc. of a terminal such as a computer, a mobile phone, a tablet, a vehicle-mounted terminal, etc., or can also be executed by a cloud server.
[0051] The technical solution of the embodiment of the present application, by obtaining the spatial information of the deployment space, the position information of each camera in the deployment space, and the video stream of each camera, determines a first target camera according to each position information and the spatial information, performs speaker recognition on each video stream to obtain a speaker recognition result, and determines a second target camera according to the speaker recognition result, where the second target camera is used to capture the speaker, determines the main screen area and the picture-in-picture area of the display screen, displays the video stream of the first target camera in the picture-in-picture area, and displays the video stream of the second target camera in the main screen area, automatically recognizes the speaker and displays it in the main screen area, facilitating intuitive and clear observation of the speaker, and at the same time displaying the video stream of the first target camera with the overall visual effect in the picture-in-picture area, which is beneficial for the speaker to grasp the overall atmosphere and adjust the rhythm and content of the speech.
[0052] In one embodiment, a number of cameras are pre-deployed in a deployment space, which includes but is not limited to meeting rooms, auditoriums, classrooms, etc. At the same time, a space model is constructed based on the deployment space to obtain the space information of the deployment space and the position information (such as coordinates) of each camera; each camera can obtain a video stream, and each video stream contains image sequence data. Each camera has a corresponding number, thereby forming a position information sequence camera_position[i], including the position information of the i-th camera, and a video stream sequence stream_active[i], including the video stream of the i-th camera.
[0053] In one embodiment, the space information includes at least one of the central position (coordinates) of the entire deployment space and specific position tags. The specific position tags may include, but are not limited to, the position coordinates of specific targets such as the podium, the hosting platform, or the center position of the conference table.
[0054] In one embodiment, step S200 includes step S210 or S220:
[0055] S210. When the space information includes the central position, according to each position information and the central position, determine the first distance between each camera and the central position. The camera with the smallest first distance among the movable cameras is used as the first target camera, or the camera with the largest first distance is used as the first target camera, or the camera with the smallest first distance is used as the first target camera.
[0056] Optionally, when the space information includes the central position, according to each position information and the central position, determine the first distance between each camera and the central position. The camera with the smallest first distance among the movable cameras is used as the first target camera. It should be noted that the movement of the camera can be up and down or left and right, so that the camera set at or near the central position moves, thereby obtaining the overall visual effect near the central position of the deployment space and comprehensively understanding the situation near the central position of the deployment space.
[0057] In some embodiments, the camera with the smallest first distance is used as the first target camera, and the camera near the central position of the deployment space is directly used as the first target camera. When the camera cannot move, try to obtain the overall visual effect near the central position of the deployment space. At the same time, the camera closest to the central position can be set as a panoramic camera in advance, and the situation near the central position of the deployment space can also be comprehensively understood. In addition, the camera with the largest first distance can also be used as the first target camera. Since it is the farthest from the central position, the viewing angle is larger at this time, so the overall visual effect can also be obtained to achieve panoramic monitoring and grasp the overall situation of the deployment space.
[0058] S220. When the spatial information includes specific location tags, determine the second distance between each camera and the specific location tag according to each location information and the specific location tag, and use the camera with the smallest second distance as the first target camera.
[0059] Optionally, when the spatial information includes specific location tags, determine the second distance between each camera and the specific location tag according to each location information and the specific location tag, and use the camera with the smallest second distance as the first target camera, so as to obtain the overall visual effect from the perspective of the specific location tag. For example, if the specific location tag is the position coordinates of the podium, the first target camera determined at this time can obtain the overall visual effect from the perspective of the host, which is convenient for understanding the overall situation below the stage.
[0060] In one implementation, in step S300, speaker recognition is performed on each video stream to obtain a speaker recognition result, including steps S310 - S320:
[0061] S310. Perform face recognition on each video stream to determine whether there is a face in each video stream.
[0062] Optionally, face detection algorithms can be used to perform face recognition on each video stream in real time, so as to determine whether there is a face in each video stream.
[0063] S320. Perform pose recognition on the video streams with faces to determine whether there is a speaker with a speaking pose, and obtain a speaker recognition result.
[0064] Optionally, if there are faces in some video streams, further perform pose recognition on these video streams through pose recognition algorithms at this time, so as to determine whether the people in the video streams conform to the speaking pose. For example, determine whether it is a speaking pose by judging whether the mouth is open and whether there are changes in the lip shapes of adjacent frames of the video stream. Therefore, finally, a speaker recognition result with a speaker conforming to the speaking pose or a speaker recognition result without a speaker conforming to the speaking pose can be obtained.
[0065] In one implementation, in step S300, according to the speaker recognition result, determine the second target camera, including steps S330 - S340:
[0066] S330. When the speaker recognition result is that there is a speaker with a speaking pose, determine the camera corresponding to the video stream in which the speaker is recognized as the second target camera.
[0067] Optionally, when the speaker recognition result indicates that there is a speaker meeting the speaking posture, determine the camera corresponding to the video stream of the recognized speaker as the second target camera. Since the presence of a speaker meeting the speaking posture indicates that the person captured by this camera is the speaker, at this time, this camera is used as the second target camera. Therefore, the second target camera is used to capture the video stream of the speaker.
[0068] S340. When the speaker recognition result indicates that there is no speaker meeting the speaking posture, use the camera corresponding to the video stream with the default face as the second target camera, or determine the second target camera from the cameras corresponding to the video streams with faces.
[0069] Optionally, the second target camera is preferentially used to capture the video stream of the speaker. When the speaker recognition result indicates that there is no speaker meeting the speaking posture, that is, there is no speaker at present. Since the video stream of the second target camera needs to be displayed on the display screen later, to ensure that the displayed content is not blank, at this time, a default face can be set, such as the host or the main person (leader, guest). Use the camera corresponding to the video stream with the default face as the second target camera, or, in the case where no default face is set, determine the second target camera from the cameras corresponding to the video streams with faces. For example, the camera corresponding to the video stream of the speaker determined last time can be used as the second target camera, or the camera corresponding to the video stream with the largest number of faces can be used as the second target camera.
[0070] In one implementation manner, determining the main screen area and the picture-in-picture area of the display screen in step S400 includes steps S410 - S430:
[0071] S410. Obtain the screen resolution of the display screen to determine the main screen area, and determine whether there is a preset picture-in-picture ratio.
[0072] Optionally, obtain the screen resolution of the display screen, and then the main screen area can be determined as the entire display area of the display screen; then, the system determines whether there is a preset picture-in-picture ratio and a preset display position set by the user in advance. If so, the size and specific display position of the picture-in-picture area can be determined. If there is no preset display position, then the picture-in-picture area is displayed at the default position, such as the upper left corner or the lower right corner.
[0073] S420. When there is a preset picture-in-picture ratio, determine the picture-in-picture area according to the product of the preset picture-in-picture ratio and the screen resolution.
[0074] Optionally, when there is a preset picture-in-picture ratio, such as 20%, the picture-in-picture area is determined according to the product of the preset picture-in-picture ratio and the screen resolution. At this time, the size of the picture-in-picture area is 20% of the screen size of the display screen. It should be noted that the picture-in-picture area is the area displayed in the form of picture-in-picture on the main screen area.
[0075] S430. When there is no preset picture-in-picture ratio, based on the default picture-in-picture ratio or the picture-in-picture ratio input after prompting, the picture-in-picture area is determined according to the product of the picture-in-picture ratio and the screen resolution.
[0076] Optionally, when there is no preset picture-in-picture ratio, based on the default picture-in-picture ratio or prompt the user to input the picture-in-picture ratio. Then, similarly, the picture-in-picture area is determined according to the product of the picture-in-picture ratio and the screen resolution.
[0077] In one implementation, in step S400, after determining the main screen area and the picture-in-picture area of the display screen, at this time, the video streams of the corresponding first target camera and the second target camera are rendered. The video stream of the first target camera is displayed in the picture-in-picture area, so that the picture-in-picture area displays the video stream of the first target camera with the overall visual effect, which is beneficial for speakers, hosts, etc. to grasp the overall atmosphere and adjust the rhythm and content of the speech. And the video stream of the second target camera is displayed in the main screen area, that is, the video stream of the speaker is displayed at a larger display ratio, and the speaker can be observed intuitively and clearly.
[0078] In one implementation, the intelligent picture-in-picture display method for multi-video stream analysis in the embodiments of the present application may further include steps S101-S102:
[0079] S101. When the video stream of the second target camera identifies the speaker, the ID continuity of the speaker is tracked through a tracking algorithm.
[0080] Optionally, when the speaker is identified each time, the system records (for example, constructs) the ID (speaker_stream_id) of the speaker, and marks the corresponding video stream with the speaker_stream_id. Then, when the video stream of the second target camera identifies the speaker, in some cases, the speaker may move and go out of the shooting range of the current second target camera. At this time, the ID continuity of the speaker is tracked through a tracking algorithm (such as optical flow method, target tracking) to determine whether the video stream captured by the current second target camera still captures the speaker. It should be noted that the embodiments of the present application track through ID tracking, and do not need to repeatedly identify the face, only need to identify whether the posture conforms to the speaking posture (if not, switch and play the video stream of the new speaker) and whether the speaker with the current ID is still captured.
[0081] S102. When the ID continuity representation of the speaker is inconsistent, match the ID of the speaker with the IDs in the video streams of other cameras except the current second target camera, and update the camera corresponding to the successfully matched video stream as the new second target camera.
[0082] Optionally, when the ID continuity representation of the speaker is inconsistent, it indicates that the current second target camera can no longer capture the current speaker. It is possible that the current speaker has walked out of the lens of the current second target camera and into other cameras. At this time, match the ID of the speaker with the IDs in the video streams of other cameras except the current second target camera, that is, compare the IDs, and if they are the same, the match is successful. Then, update the camera corresponding to the successfully matched video stream as the new second target camera, so as to realize the dynamic tracking and adjustment of the speaker. Among them, in the embodiments of the present application, by means of ID matching, compared with face comparison or pose comparison, it is simpler and more direct, reduces the data processing volume of the system, and has a higher accuracy rate. In addition, the video stream of the updated second target camera will replace the video stream of the old second target camera currently played on the display screen in real time, and based on the change of the second target camera, realize the dynamic change adjustment of the playback content.
[0083] In an implementation manner, when the video stream of the first target camera is displayed in the picture-in-picture area and the video stream of the second target camera is displayed in the main screen area, the intelligent picture-in-picture display method for multi-video stream analysis in the embodiments of the present application may further include steps S103-S104:
[0084] S103. In response to a click instruction from the user on the picture-in-picture area, switch the picture-in-picture area and the main screen area.
[0085] Optionally, in order to meet the actual needs of different scenarios and different users, the embodiments of the present application provide a dynamic adjustment function. The user can perform a click operation on the picture-in-picture area on the display screen to generate a click instruction. The system responds to the click instruction and switches the picture-in-picture area and the main screen area. At this time, the content in the picture-in-picture area is displayed in the main screen area, and the content in the main screen area is displayed in the picture-in-picture area.
[0086] S104. In response to a drag instruction from the user on the picture-in-picture area, move the picture-in-picture area or adjust the size of the picture-in-picture area.
[0087] Optionally, the user can, based on actual needs, for example, if the current picture-in-picture area may block some key information in the main screen area, the user can generate a drag instruction by performing a drag operation on the picture-in-picture area. In response to the drag instruction, the system moves the picture-in-picture area or adjusts the size of the picture-in-picture area to meet the needs of different users.
[0088] In one implementation, key content can also be preset, and then it is automatically recognized whether the preset key content in the main screen area is blocked by the picture-in-picture area. If so, the position of the picture-in-picture area is automatically adjusted to display the complete preset key content.
[0089] Referring to Figure 2 , a structural block diagram of an intelligent picture-in-picture display device for multi-video stream analysis according to an embodiment of the present application is shown. The device may include:
[0090] An acquisition module, configured to acquire spatial information of a deployment space, position information of each camera in the deployment space, and video streams of each camera;
[0091] A first determination module, configured to determine a first target camera according to each position information and the spatial information, where the first target camera is used to provide an overall visual effect;
[0092] A second determination module, configured to perform speaker recognition on each video stream to obtain a speaker recognition result, and determine a second target camera according to the speaker recognition result, where the second target camera is used to capture the speaker;
[0093] A display module, configured to determine a main screen area and a picture-in-picture area of a display screen, display the video stream of the first target camera in the picture-in-picture area, and display the video stream of the second target camera in the main screen area.
[0094] For the functions of the modules in the device according to the embodiment of the present application, reference may be made to the corresponding descriptions in the above method, which will not be elaborated here.
[0095] Referring to Figure 3 , a structural block diagram of an electronic device according to an embodiment of the present application is shown. The electronic device includes: a memory 310 and a processor 320. Instructions that can run on the processor 320 are stored in the memory 310. The processor 320 loads and executes the instructions to implement the intelligent picture-in-picture display method for multi-video stream analysis in the above embodiment. Among them, the number of the memory 310 and the processor 320 can be one or more.
[0096] In one embodiment, the electronic device further includes a communication interface 330 for communicating with external devices and performing data interaction and transmission. If the memory 310, the processor 320, and the communication interface 330 are implemented independently, the memory 310, the processor 320, and the communication interface 330 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 only a thick line is used to represent it in Figure 3 , but it does not mean that there is only one bus or one type of bus.
[0097] Optionally, in a specific implementation, if the memory 310, the processor 320, and the communication interface 330 are integrated on a single chip, the memory 310, the processor 320, and the communication interface 330 can communicate with each other through an internal interface.
[0098] The embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the intelligent picture-in-picture display method for multi-video stream analysis provided in the above embodiment.
[0099] The embodiment of the present application further provides a chip including a processor for calling and running instructions stored in a memory, so that a communication device installed with the chip executes the method provided in the embodiment of the present application.
[0100] The embodiment of the present application further provides a chip including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute code in the memory, and when the code is executed, the processor is configured to execute the method provided in the embodiment of the application.
[0101] It should be understood that the above-mentioned processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the advanced RISC machines (ARM) architecture.
[0102] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may also include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both a volatile and a non-volatile memory. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0103] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0104] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0105] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of these features. In the description of the present application, "a plurality of" means two or more unless otherwise specifically defined.
[0106] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed.
[0107] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices.
[0108] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.
[0109] In addition, in each embodiment of this application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0110] As mentioned above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of various changes or substitutions, and these should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. An intelligent picture-in-picture display method for multi-video stream analysis, characterized in that, Including: Obtaining spatial information of a deployment space, position information of each camera in the deployment space, and video streams of each of the cameras; Determining a first target camera according to each of the position information and the spatial information, where the first target camera is used to provide an overall visual effect; Performing speaker recognition on each of the video streams to obtain a speaker recognition result, and determining a second target camera according to the speaker recognition result, where the second target camera is used to capture a speaker; Determining a main screen area and a picture-in-picture area of a display screen, displaying the video stream of the first target camera in the picture-in-picture area, and displaying the video stream of the second target camera in the main screen area.
2. The intelligent picture-in-picture display method for multi-video stream analysis according to claim 1, wherein: The determining the first target camera according to each of the position information and the spatial information includes: When the spatial information includes a central position, determining a first distance between each camera and the central position according to each of the position information and the central position, and taking the camera with the smallest first distance among the movable cameras as the first target camera, or taking the camera with the largest first distance as the first target camera, or taking the camera with the smallest first distance as the first target camera; Or, When the spatial information includes a specific position tag, determining a second distance between each camera and the specific position tag according to each of the position information and the specific position tag, and taking the camera with the smallest second distance as the first target camera.
3. The intelligent picture-in-picture display method for multi-video stream analysis according to claim 1, characterized in that: The performing speaker recognition on each of the video streams to obtain a speaker recognition result includes: Performing face recognition on each of the video streams to determine whether there is a face in each video stream; Performing pose recognition on the video streams with faces to determine whether there is a speaker with a speaking pose, and obtaining a speaker recognition result.
4. The intelligent picture-in-picture display method for multi-video stream analysis according to claim 3, characterized in that: The determining the second target camera according to the speaker recognition result includes: When the speaker recognition result is that there is a speaker with a speaking pose, determining the camera corresponding to the video stream in which the speaker is recognized as the second target camera; When the speaker recognition result is that there is no speaker with a speaking pose, taking the camera corresponding to the video stream with a default face as the second target camera, or determining the second target camera from the cameras corresponding to the video streams with faces.
5. The intelligent picture-in-picture display method for multi-video stream analysis according to claim 4, wherein: The method further includes: When a speaker is recognized in the video stream of the second target camera, tracking the ID continuity of the speaker through a tracking algorithm, where when a speaker is recognized, the ID of the speaker is recorded; When the ID continuity of the speaker is characterized as inconsistent, matching the ID of the speaker with the IDs in the video streams of other cameras except the current second target camera, and updating the camera corresponding to the video stream with a successful match as the new second target camera.
6. The intelligent picture-in-picture display method for multi-video stream analysis according to any one of claims 1-5, characterized in that: The determining the main screen area and the picture-in-picture area of the display screen includes: Obtaining the screen resolution of the display screen to determine the main screen area, and judging whether there is a preset picture-in-picture ratio; When there is a preset picture-in-picture ratio, determine the picture-in-picture area according to the product of the preset picture-in-picture ratio and the screen resolution; When there is no preset picture-in-picture ratio, based on the default picture-in-picture ratio or the picture-in-picture ratio input after prompting, determine the picture-in-picture area according to the product of the picture-in-picture ratio and the screen resolution.
7. The intelligent picture-in-picture display method for multi-video stream analysis according to claim 6, characterized in that: The method further includes: When the video stream of the first target camera is displayed in the picture-in-picture area and the video stream of the second target camera is displayed in the main screen area, it includes at least one of the following: In response to a click instruction from the user on the picture-in-picture area, switch the picture-in-picture area and the main screen area; In response to a drag instruction from the user on the picture-in-picture area, move the picture-in-picture area or adjust the size of the picture-in-picture area.
8. An intelligent picture-in-picture display device for multi-video stream analysis, characterized in that, It includes: An acquisition module, configured to acquire the spatial information of the deployment space, the position information of each camera in the deployment space, and the video stream of each camera; A first determination module, configured to determine a first target camera according to each of the position information and the spatial information, where the first target camera is used to provide an overall visual effect; A second determination module, configured to perform speaker recognition on each of the video streams to obtain a speaker recognition result, and determine a second target camera according to the speaker recognition result, where the second target camera is used to capture the speaker; A display module, configured to determine the main screen area and the picture-in-picture area of the display screen, display the video stream of the first target camera in the picture-in-picture area, and display the video stream of the second target camera in the main screen area.
9. An electronic device, characterized in that, It includes: A processor and a memory, where instructions are stored in the memory, and the instructions are loaded and executed by the processor to implement the method according to any one of claims 1-7.
10. A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it implements the method according to any one of claims 1-7.