Data processing method, device and system
By establishing automatic and triggered connection relationships in a large conference room system and outputting images based on the target mode, the cost and difficulty problems brought by multiple cameras are solved, and efficient image processing and user experience improvement are achieved.
Patent Information
- Application Number
- CN202510467488.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-08
AI Technical Summary
The use of multiple cameras in large conference room systems leads to increased cost and increased deployment difficulty.
Image data is automatically acquired by establishing a first connection relationship based on historical configuration parameters on the host device, and the second device triggers the establishment of a second connection relationship to acquire image data during the operation stage, and the object is determined in combination with the image data and output the image based on the target mode, so that the images of different objects are located in different output areas.
It reduces the cost and difficulty of deploying additional cameras for host devices, improves the efficiency of image data acquisition and processing, and enhances the user experience.
Smart Images

Figure CN120455619A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronic technology, and in particular to a data processing method, device, and system. Background Art
[0002] Large conference room systems typically support multiple camera scenarios, enabling video output options, including single-camera selection and multi-camera video layout. However, these scenarios require the use of conference machines with dedicated cameras, increasing costs and deployment complexity. Summary of the Invention
[0003] In view of this, embodiments of the present application provide at least one data processing method.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] In a first aspect, an embodiment of the present application provides a data processing method, the data processing method comprising:
[0006] obtaining image data from at least one first device based on a first connection relationship, wherein the first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device;
[0007] obtaining image data from at least one second device based on a second connection relationship, where the second connection relationship is triggered and established by the second device during the current operation phase of the host device;
[0008] determining a plurality of objects and a target image corresponding to each object based on the image data;
[0009] The target images are output based on the target mode so that target images corresponding to different objects are located in different output areas.
[0010] In a second aspect, an embodiment of the present application provides a data processing method, the data processing method comprising:
[0011] In response to the target instruction, establishing a second connection relationship with the host device;
[0012] sending image data to the host device based on the second connection relationship;
[0013] The host device is further capable of obtaining image data of the first device based on a first connection relationship, where the first connection relationship is automatically established based on historical configuration parameters for the first device after the host device is started.
[0014] In a third aspect, an embodiment of the present application provides a data processing device, the device comprising:
[0015] a first obtaining module, configured to obtain image data from at least one first device based on a first connection relationship, wherein the first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device;
[0016] a second obtaining module, configured to obtain image data from at least one second device based on a second connection relationship, wherein the second connection relationship is triggered and established by the second device during the current operation phase of the host device;
[0017] a determination module, configured to determine a plurality of objects and a target image corresponding to each object based on the image data;
[0018] The output module is used to output each target image based on the target mode, so that the target images corresponding to different objects are located in different output areas.
[0019] In a fourth aspect, an embodiment of the present application provides a data processing system, the system comprising a host device, at least one first device, and at least one second device; wherein,
[0020] The first device and the second device are used to collect image data;
[0021] The host device is configured to obtain image data from the first device and the second device based on a first connection relationship and a second connection relationship, respectively; the first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device, and the second connection relationship is triggered by the second device during the current operation of the host device;
[0022] The host device is further configured to determine multiple objects and target images corresponding to each object based on the image data; and output each target image based on a target mode so that target images corresponding to different objects are located in different output areas.
[0023] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.
[0025] Figure 1a A schematic diagram of the structure of a conference system in related technology;
[0026] Figure 1b Schematic diagram 1 of an implementation flow of a data processing method provided in an embodiment of the present application;
[0027] Figure 2 Scenario diagram 1 of a data processing method provided in an embodiment of the present application;
[0028] Figure 3 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 2 ;
[0029] Figure 4 A scenario diagram of a data processing method provided in an embodiment of the present application Figure 2 ;
[0030] Figure 5 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 3 ;
[0031] Figure 6 A scenario diagram of a data processing method provided in an embodiment of the present application Figure 3 ;
[0032] Figure 7 A scenario diagram of a data processing method provided in an embodiment of the present application Figure 4 ;
[0033] Figure 8 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 4 ;
[0034] Figure 9 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;
[0035] Figure 10 A schematic diagram of the structure of a data processing system provided in an embodiment of the present application;
[0036] Figure 11 This is a schematic diagram of a hardware entity of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0038] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0039] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing this application only and are not intended to limit this application.
[0041] Large conference room systems typically support multiple camera scenarios, enabling video output options, including single camera selection and multi-camera video layout. These scenarios require the use of conference machines with dedicated cameras, but the additional cameras increase costs and deployment complexity.
[0042] For example, Figure 1a As shown, a conference room system 11 in the related art includes multiple cameras 111 and multiple display devices 112; a user terminal 12 is connected to the conference room system 11. The multiple cameras 111 of the conference room system 11 are used to capture multi-angle images of the user, and the multiple display devices 112 are used to display the multi-angle images of the user. As can be seen, in the related art, in order to capture multi-angle images of the user, multiple cameras need to be deployed in the conference room system, which increases costs and deployment difficulties.
[0043] To address the above technical issues, embodiments of the present application provide a data processing method that can be applied to an electronic device, such as a host device having a processor, or a conference machine in a conference system. In some embodiments, the electronic device can also be an audio-video control box (AV box) in the conference system.
[0044] Figure 1b A data processing method according to an embodiment of the present invention is provided as a flowchart. Figure 1b As shown, the data processing method can be implemented through steps S101 to S104:
[0045] Step S101 : obtaining image data from at least one first device based on a first connection relationship, where the first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device.
[0046] Step S102: Obtain image data from at least one second device based on a second connection relationship, where the second connection relationship is triggered and established by the second device during the current operation phase of the host device.
[0047] In the embodiments of the present application, the first device may be a camera corresponding to the host device. The second device may be a user's terminal device, which may include but is not limited to a smartphone, tablet computer, wearable device, personal computer (PC), netbook, etc. In some embodiments, the second device may also refer only to a built-in or external camera of the terminal device.
[0048] For example, in a conference scenario, the first device and the second device are different devices with image capture capabilities in the same conference room. The first device may be a camera in the conference system, which captures image data of participants. The second device may be a terminal device of the participant, which captures image data of participants using a built-in or external camera of the terminal device.
[0049] It is understandable that in a conference scenario, because the first device and the host device are both hardware devices of the conference system, under normal circumstances, the AV device (Audio / Video device) in the conference room will establish a connection relationship with the host device when it is first deployed, and can directly connect to the host device the next time it is used. Therefore, the first connection relationship between the first device and the host device is automatically established after the host device is started based on the historical configuration parameters for the first device. Therefore, when the host device and the first device are started, the first connection relationship between the two is automatically established, and the image data collected by the first device can be directly transmitted to the host device based on the first connection relationship.
[0050] For the second device, because it is a device belonging to an offline participant, a second connection relationship between the second device and the host device cannot be automatically established after the second device and the host device are started. During the host device's current operation phase, the second device must first trigger the establishment of a second connection relationship with the host device. Then, after the second device captures the participant's image data, it sends the image data to the host device. In other words, the first device can directly send the captured image data to the host device, while the second device first establishes a communication connection with the host device and then sends the captured image data to the host device. Therefore, the method for obtaining image data for the first device is different from the method for obtaining image data for the second device.
[0051] In an embodiment of the present application, the second device can join the network where the host device is located, thereby triggering the establishment of a communication connection between the two.
[0052] In some embodiments, the second device may also connect to the host device via Bluetooth, thereby triggering the establishment of a communication connection between the two.
[0053] In some embodiments, the second device may also first send the image data to the cloud corresponding to the host device, and then the cloud sends the image data to the host device.
[0054] Step S103 : determining a plurality of objects and a target image corresponding to each object based on the image data.
[0055] Here, the object may be an object in the image data. For example, in a conference scenario, the object may be a participant. The target image may be an image that can show the front face of the object.
[0056] In an embodiment of the present application, image data from different sources may be classified first to obtain image data sets corresponding to different objects; and then the corresponding target image of the object may be determined in the image data sets.
[0057] It is understandable that because the first and second devices capture images from different perspectives, the target image of an object may be image data captured by the first device or the second device. Therefore, it is necessary to first classify multiple image data of the same object together and then determine the corresponding target image of the object in the image data set.
[0058] In some embodiments, image data from different sources may be classified based on the anthropomorphic features of the subject images in the image data to obtain image data sets corresponding to different subjects. For example, the anthropomorphic features may include at least one of the following: skeletal proportions (e.g., arm span, head-to-body ratio, joint spacing), body type (height, weight, body type), clothing color / texture, hairstyle, shoe style, backpack, hat, accessories, etc.
[0059] In some embodiments, the face deflection angle corresponding to the image data may be determined first; and the image data whose face deflection angle satisfies the deflection condition may be determined as the target image corresponding to the object.
[0060] In some embodiments, when there are multiple objects in the image data, the image data may be segmented to obtain image sub-data, and then the multiple image sub-data may be classified.
[0061] It is understandable that in a conference scenario, the image acquisition range of the first device is often relatively large and can capture image data of multiple participants, so it is necessary to segment the image data captured by the first device. The second device may also capture image data of multiple participants.
[0062] Step S104 : outputting the target images based on the target mode, so that target images corresponding to different objects are located in different output areas.
[0063] Here, the target mode may be an output mode for the image data. In some embodiments, the target mode may be a mode for simultaneously displaying multiple image data, such as a multi-video screen arrangement display mode. In other embodiments, the target mode may also be a mode for sequentially displaying multiple image data, such as an image carousel mode.
[0064] For example, Figure 2 As shown, four target images 51 are output simultaneously according to a multi-video screen arrangement display mode.
[0065] In some embodiments, different target images may be fused to obtain a fused image corresponding to the target pattern, and then the fused image may be output.
[0066] It is understood that the first device and the second device capture video stream data, so when performing the fusion process, the frame image captured by the first device needs to be spliced with the frame image corresponding to the second device. Specifically, when splicing the frame images, the frame image captured by the first device and the frame image corresponding to the second device at each image capture time point need to be spliced.
[0067] In some embodiments, after determining the target image of each object, an output region for each target image may be determined, and then the corresponding target image may be output in the output region. In other words, in the embodiments of the present application, there is no need to perform fusion processing on multiple target images; the corresponding target image may be output in the corresponding output region.
[0068] In the embodiments of the present application, image data captured by different devices can be acquired, and then multiple objects and target images corresponding to each object can be determined based on the image data. Each target image can be output based on a target mode, so that target images corresponding to different objects are located in different output areas. In this way, the image acquisition capabilities of different types of devices can be utilized to output target images of various objects. This allows the acquisition of target images of objects from multi-angle image data even when a small number of first devices are deployed, thereby reducing the cost and difficulty associated with deploying the first device on the host device.
[0069] In some embodiments, as Figure 3 As shown, the above step S102 can be implemented through steps S301 and S302:
[0070] Step S301 : Compare at least two registered images of registered objects in the database with the image data of the multiple devices to obtain at least two image data of each of the objects; different image data of each of the objects corresponds to different image acquisition angles.
[0071] In an embodiment of the present application, the host device can obtain at least two registered images corresponding to different registered objects from its own database, and then compare the registered images with each image data, thereby classifying the at least two image data according to different objects.
[0072] It is understandable that the image acquisition range of the first device is different from that of the second device, so the perspectives of the images captured by different devices are different. For the same object, it may be captured by both the first device and the second device. At this time, the image data of multiple devices will contain multi-perspective images of the same object. Therefore, in order to classify the multiple image data according to the object, it is necessary to obtain the pre-registered registration images of multiple registered objects from the database of the host device and compare them with the image data. Because the perspective of the registration image of the registered object may be different from the perspective of the image data, it is necessary to obtain multi-angle registration images of each registered object and compare them with the image data, so as to improve the accuracy of image comparison.
[0073] In an embodiment of the present application, the registrant can log in to the registration webpage or application corresponding to the host device through its own terminal device in advance, and then upload its own registered images from multiple perspectives to the host device.
[0074] In an embodiment of the present application, after obtaining the registered image, the human body features of the registered image and the human body features of the image data can be collected first, and then the human body features between the two images can be compared to achieve object classification of the image data.
[0075] Step S302 : determining a target image corresponding to each object based on at least two image data of each object.
[0076] Here, the target image corresponding to the object may be an image with the best acquisition angle of the object among the at least two image data, or an image with the highest similarity to the registered image.
[0077] In an embodiment of the present application, for each object, the face angles of at least two image data of the object can be determined, and the image data that meets the face deflection angle threshold is determined as the target image.
[0078] In this embodiment of the present application, the number of facial key points in each image data can be determined first, and then first image data whose number of facial key points meets a threshold can be obtained. The facial deflection angle in the first image data can then be determined, and the first image data whose facial deflection angle meets a preset angle can be determined as the target image. For example, the threshold can be 68 or 109.
[0079] It is understandable that when the face is turned sideways, it is impossible to collect all the feature points of the face, so the image data can be filtered by the number of facial key points, and then by calculating the facial deflection angle of the face, the image data that meets the facial deflection angle threshold can be determined.
[0080] In some embodiments, the above step S302 may also be implemented through steps S3021 and S3022:
[0081] Step S3021 : determining the similarity between a target registered image in the at least two registered images of the registered object and the at least two image data of the object.
[0082] The target registration image can be the frontal face image uploaded by the participant during registration. Alternatively, it can be the participant's preferred facial perspective. For example, some people find facial contours more harmonious when viewed 20 degrees from the left, while others find them more harmonious when viewed 20 degrees from the right. This allows for user-defined optimal viewing angles, always presenting the participant's optimal facial perspective.
[0083] Exemplarily, at least one first device and at least one second device include electronic device 1 located to the left of participant A and electronic device 2 located to the right of participant A. That is, electronic device 1 captures image data 1 of the left side of participant A, and electronic device 2 captures image data 2 of the right side of participant A. When the target registration image uploaded by participant A during the registration phase is left side image data, it can be determined that the similarity between image data 1 and the left side image data is greater than the similarity between image data 2 and the left side image data. Therefore, image data 1 can be determined as the target image corresponding to participant A, that is, the host device can output image data 1 captured by electronic device 1.
[0084] In this embodiment of the present application, during the registration phase, the subject is required to upload multiple registration images of themselves at different acquisition angles, including a frontal perspective. Furthermore, the host device tags the different registration images with their respective angles during upload. Therefore, the target registration image of the subject can be obtained through the angle tags, and then the target registration image can be compared with the image data for similarity.
[0085] Exemplarily, the similarity comparison method may include at least one of the following: pixel difference method, histogram statistics method, hash algorithm, deep learning, etc.
[0086] Step S3022: Determine the image data that meets the similarity threshold as the target image corresponding to the object.
[0087] In the embodiment of the present application, image data that meets a similarity threshold may be determined as a target image corresponding to the object. For example, the similarity threshold may be 80%.
[0088] In an embodiment of the present application, at least two registered images of registered objects in a database are compared with image data from multiple devices to obtain at least two image data sets for each object. Different image data sets for each object correspond to different image acquisition angles. Based on the at least two image data sets for each object, a target image corresponding to each object is determined. In this way, image data from multiple devices is classified using the registered images of the registered objects, allowing the currently acquired image data to be associated with the registered objects. This allows the registered object to determine the target image it considers optimal from the image data acquired from multiple image acquisition angles, thereby improving the user experience.
[0089] In some embodiments, when the similarity between the currently output target image and the target registered image does not meet the similarity threshold, it is necessary to re-determine the similarity between the target registered image in at least two registered images of the registered object and at least two image data of the object, and determine the image data that meets the similarity threshold as the target image corresponding to the object, and then output the re-determined target image.
[0090] In some embodiments, the above method may also be implemented through step S11, and correspondingly, the above step S104 may be implemented through step S12:
[0091] Step S11 : comparing the registered images of the registered objects in the database with the target images corresponding to the objects to obtain the object attribute information corresponding to the objects.
[0092] Exemplarily, the object attribute information may be the name information of the registered object, or may be the job title information of the registered object.
[0093] In the embodiment of the present application, the similarity between the registered image of the registered object in the database and the target image corresponding to each object can be determined to obtain the registered image corresponding to each object. Then, based on the correspondence between the registered image and the object attribute information, the object attribute information of each object can be determined.
[0094] It is understood that during the registration phase, the user must upload not only their own registration images from different viewing angles to the host device, but also their own object attribute information. This allows the host device's database to store the corresponding relationships between different registration images and the object attribute information of different registration objects.
[0095] Step S12: outputting each target image and object attribute information corresponding to each target image based on the target mode.
[0096] In the embodiment of the present application, in the process of outputting the target image, object attribute information corresponding to the target image can be output together.
[0097] In the embodiments of the present application, by comparing the registered images of registered objects in the database with the target images corresponding to each object, object attribute information corresponding to each object can be obtained. Then, based on the target mode, each target image and the object attribute information corresponding to each target image are output. This allows the object attribute information of the objects in the target image to be output simultaneously with the target image, thereby improving the user experience.
[0098] In some embodiments, the above method may also be implemented through steps S21 and S22:
[0099] Step S21, in response to establishing a second connection relationship with the second device, receiving at least two registration images of the registration object and object attribute information of the registration object; the registration image and the object attribute information are sent by the second device or the terminal device to which the second device belongs based on the registration page, and the registration page is a page displayed on the second device or the terminal device to which the second device belongs.
[0100] In an embodiment of the present application, when a second connection relationship is established between the second device and the host device, the second device or the terminal device to which the second device belongs can display a registration page, and the registered object can upload its own registration image and object attribute information to the host device through the registration page, so that the host device receives at least two registration images of the registered object and the object attribute information of the registered object.
[0101] It is understood that after the second device completes the establishment of the second connection relationship with the host device, the second device or the terminal device to which the second device belongs can display a registration interface, thereby allowing the registered object to upload its own registration image and object attribute information. In other words, the second device or the terminal device to which the second device belongs needs to establish the second connection relationship first, and then the second device or the terminal device to which the second device belongs displays the registration interface.
[0102] Step S22: storing the corresponding relationship between the registered image and the object attribute information in the database.
[0103] In an embodiment of the present application, after the host device receives the registration image and object attribute information, a correspondence between different registration images and different object attribute information can be established based on the device attribute information of the second device that sends the registration image and object attribute information or the terminal device to which the second device belongs, or the IP address.
[0104] In this embodiment of the present application, in response to establishing a second connection relationship with a second device, at least two registered images of a registered object and object attribute information of the registered object are received, and then the correspondence between the registered images and the object attribute information is stored in the database. In this way, the correspondence between different registered images and different object attribute information can be stored in the database of the host device, thereby identifying corresponding object attribute information for image data collected by different devices.
[0105] In some embodiments, the above method may also be implemented through step S31:
[0106] Step S31, outputting each of the target images and image data of the first device based on the target mode;
[0107] The image acquisition range of the first device is larger than the image acquisition range of the second device.
[0108] In the embodiment of the present application, the image data collected by the first device can be output while each target image is output according to the target mode.
[0109] It should be noted that the target image can be part of the image data of the first device. For example, in a conference scenario, the first device is a wide-angle camera of the conference system, which can capture image data of all participants. When a participant turns toward the first device, the first device can capture the target image of that participant.
[0110] For example, Figure 4 As shown, while the host device outputs four target images 51 , it also outputs image data 61 captured by the first device.
[0111] In this embodiment of the present application, the host device can simultaneously output image data of the first device with a larger image acquisition range when outputting different target images based on the target mode. This allows the subject to see image data from different perspectives, thereby improving the subject's experience.
[0112] In some embodiments, the predicted position information of a speaking subject among multiple subjects may be determined based on the image data captured by the first device. The speaking subject may then be determined based on image data captured by at least one second device corresponding to the predicted position information, and the target object corresponding to the speaking subject may be labeled.
[0113] It is understandable that the image data collected by the first device is video data, and the video data includes audio data, so the predicted position information of the speaking object can be determined through the audio data.
[0114] In an embodiment of the present application, determining the speaking object based on image data collected by at least one second device corresponding to the predicted position information may include: extracting a mouth image from the image data collected by the second device; and determining whether the object in the image data is the speaking object based on the mouth image.
[0115] In some embodiments, an embodiment of the present application provides a data processing method, which can be applied to a second device, or a terminal device to which the second device belongs.
[0116] Figure 5 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 3 ,like Figure 5 As shown, the data processing method can be implemented through steps S501 and S502:
[0117] Step S501: Establishing a second connection relationship with a host device in response to a target instruction.
[0118] Step S502: Send image data to the host device based on the second connection relationship.
[0119] The host device is further capable of obtaining image data of the first device based on a first connection relationship, where the first connection relationship is automatically established based on historical configuration parameters for the first device after the host device is started.
[0120] In an embodiment of the present application, the host device can also determine multiple objects and target images corresponding to each object based on the image data collected by the first device and the image data collected by the second device, and output each target image based on the target mode so that the target images corresponding to different objects are located in different output areas.
[0121] In some embodiments, the above method may also be implemented by at least one of step S41 and step S42:
[0122] Step S41 , running a first application to display a registration page, obtaining address information input based on the registration page, and establishing a connection with the host device based on the address information; the address information is the network address of the host device.
[0123] Exemplarily, the first application may be a browser application.
[0124] In an embodiment of the present application, the second device or the terminal device to which the second device belongs can, in response to a startup operation of the first application, run the first application to display a registration page. The registration page includes an address entry area that can receive address information entered by the registrant. The second device or the terminal device to which the second device belongs can establish a connection with the host device based on the address information, thereby uploading image data of the object corresponding to the second device to the host device.
[0125] It is understood that if the current scenario is a conference scenario and the second device is a participant's terminal device, after the second device establishes a connection with the host device, the second device's camera can simply serve as a remote external camera for the host device to capture participant image data. In other words, even if the second device does not join the online meeting established by the host device, the second device can still transmit the participant's image data to the host device.
[0126] Step S42 , running the second application to enter the target session, so that the cloud device that created the target session sends the image data of the second device to the host device; the host device is in the target session.
[0127] In an embodiment of the present application, the second device or the terminal device to which the second device belongs can, in response to a startup operation of the second application, run the second application to enter a target session. Through the target session, image data captured by the second device can be sent to the server of the second application, and then the server of the second application can send the image data captured by the second device to the host device in the above target session.
[0128] It can be understood that when the current scene is a conference scene and the second device is a terminal device of a participant, when the second device enters the target session by running the second application, the second device enters the online meeting, and the second device sends the image data of the participant corresponding to the second device to the host device.
[0129] Assume that there are offline participants 1 and 2, and online participants 3 and 4:
[0130] For step S41, the first device of the conference machine can collect image data of participant 1 and participant 2. The second devices of offline participant 1 and participant 2 are connected to the host device by running the first application, and the image data collected by their own devices are sent to the host device. After processing, the host device outputs the image data on the large screen of the conference room according to the target mode. Online participants 3 and participant 4 join the meeting remotely through their respective second devices, and can see the same output content as the large screen on their respective second devices. In this scenario, offline participants 1 and participant 2 only serve as extended collection devices of the host device.
[0131] For step S42: the first device of the conference machine can collect image data of participant 1 and participant 2, and the second device of each offline participant 1 and participant 2 is connected to the host device by running the second application, and sends the image data collected by its own device to the server of the second application. After processing, the server of the second application synchronously sends it to participants 1 and participant 2, online participants 3 and participant 4, and the conference machine, so that all the above devices have the same output content.
[0132] The following describes the application of the data processing method provided in the embodiment of the present application in a practical scenario:
[0133] Large conference room systems typically support multiple camera scenarios, enabling video output options, including single camera selection and multi-camera video layout. These scenarios require the use of conference machines with dedicated cameras, but the additional cameras increase costs and deployment complexity.
[0134] In order to solve the technical problems in the related art, the embodiment of the present application provides a data processing method. The execution subject of the method can be a conference machine or an audio-video control box (AV Box). The data processing method can be implemented through steps S1 to S5:
[0135] Step S1: Find all video streams containing portraits from all video streams.
[0136] In the embodiment of the present application, all video streams may include electronic device video streams captured by the camera of the electronic device and conference machine video streams captured by the camera of the conference machine.
[0137] Step S2: classify all video streams containing human portraits according to different human objects based on the human portrait features.
[0138] In an embodiment of the present application, character features may include: skeletal proportions (e.g., arm span length, head-to-body ratio, joint spacing), body shape (height, weight, and body shape), clothing color / texture, hairstyle, shoe style, backpack, hat, and accessories.
[0139] Step S3: determining a target video stream from the video streams corresponding to each person object.
[0140] In the embodiment of the present application, the target video stream can be the video stream with the best facial angle among the video streams corresponding to the human object.
[0141] Step S4: Based on the registration information, determine the names and positions of different character objects, and display the names and positions on the corresponding target video stream.
[0142] For example, Figure 6 As shown, the user can upload his / her own image, name, and position on the registration webpage 40 through his / her own terminal device. Among them, the user's image can be the user's image under different perspectives, such as Figure 6 As shown, the user images are displayed from three different perspectives. The terminal device can be a smartphone, tablet, or wearable device running different operating systems, including Android and iOS (iPhone Operating System). The terminal device can also be a personal computer or laptop running different operating systems, including Windows, Linux, and Macintosh Operating Systems.
[0143] Step S5, identifying whether it is a speaker through mouth movements, marking it on the target video stream, and then outputting the corresponding video stream according to the settings of the current mode.
[0144] For example, Figure 7 As shown, in the current conference scenario, there are four participants. The terminal device 21 corresponding to each participant captures image 2111 and sends image 2111 to the conference machine 31. The conference machine camera 311 can capture image 3111 and send image 3111 to the conference machine 31. The conference machine 31 can process image 3111 and image 2111 to output the corresponding video streams.
[0145] Below through Figure 8 To describe the above data processing method, such as Figure 8 As shown:
[0146] Terminal devices 21 through 2n upload their user registration information to conference machine 31 in advance, which stores it as user information. When a conference begins, terminal cameras 211 through 21n, respectively, corresponding to terminal devices 21 through 2n, capture image data. Conference machine camera 311 also captures image data and sends it to conference machine 31. Conference machine 31 performs the following steps:
[0147] Step S21: human body detection.
[0148] Step S22, re-identification; when a user exists in the image data, re-identification is performed, that is, image data of the same user are classified together based on pre-stored user information.
[0149] Step S23: speech detection, determining whether there is a speaker in the image data.
[0150] Step S24 , screening the image data, that is, determining the image data with the best angle of the user's face in the image data.
[0151] Step S25: Post-processing: In the embodiment of the present application, the image data may be processed based on the current data output mode of the conference machine.
[0152] Step S26: Output to the camera application.
[0153] Figure 9 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application is shown in FIG. Figure 9 As shown, the data processing device 900 includes a first obtaining module 901, a second obtaining module 902, a determining module 903 and an output module 904; wherein,
[0154] A first obtaining module 901 is configured to obtain image data from at least one first device based on a first connection relationship, where the first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device;
[0155] A second obtaining module 902 is configured to obtain image data from at least one second device based on a second connection relationship, where the second connection relationship is triggered and established by the second device during the current operation phase of the host device;
[0156] A determination module 903 is configured to determine a plurality of objects and a target image corresponding to each object based on the image data;
[0157] The output module 904 is configured to output each target image based on the target mode, so that target images corresponding to different objects are located in different output areas.
[0158] Figure 10 A schematic diagram of the structure of a data processing system provided in an embodiment of the present application is shown in FIG. Figure 10 As shown, the data processing system 1000 includes a host device 1001, at least one first device 1002 and at least one second device 1003; wherein,
[0159] The first device 1002 and the second device 1003 are used to collect image data;
[0160] The host device 1001 is configured to obtain image data from the first device and the second device based on a first connection relationship and a second connection relationship, respectively; the first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device, and the second connection relationship is triggered by the second device during the current operation of the host device;
[0161] The host device 1001 is further configured to determine multiple objects and target images corresponding to each object based on the image data; and output each target image based on a target mode so that target images corresponding to different objects are located in different output areas.
[0162] In some embodiments, the host device 1001 is further used to compare at least two registered images of registered objects in the database with the image data of the multiple devices to obtain at least two image data of each of the objects; different image data of each of the objects correspond to different image acquisition angles; based on the at least two image data of each of the objects, determine the target image corresponding to each object.
[0163] In some embodiments, the host device 1001 is further used to determine the similarity between a target registered image in at least two registered images of the registered object and at least two image data of the object; and determine the image data that meets the similarity threshold as the target image corresponding to the object.
[0164] In some embodiments, the host device 1001 is also used to compare the registered images of the registered objects in the database with the target images corresponding to the objects to obtain object attribute information corresponding to the objects; based on the target mode, output each target image and the object attribute information corresponding to each target image.
[0165] In some embodiments, the host device 1001 is also used to receive at least two registration images of the registration object and object attribute information of the registration object in response to establishing a second connection relationship with the second device; the registration image and the object attribute information are sent by the second device or the terminal device to which the second device belongs based on a registration page, and the registration page is a page displayed on the second device or the terminal device to which the second device belongs; the correspondence between the registration image and the object attribute information is stored in the database.
[0166] In some embodiments, the host device 1001 is further configured to output each of the target images and image data of the first device based on the target mode; wherein the image acquisition range of the first device is larger than the image acquisition range of the second device.
[0167] In some embodiments, the second device 1003 is also used to establish a second connection relationship with the host device in response to a target instruction; based on the second connection relationship, send image data to the host device; the host device can also obtain the image data of the first device based on the first connection relationship, and the first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device.
[0168] In some embodiments, the second device 1003 is also used to run the first application to display a registration page, obtain address information input based on the registration page, and establish a connection with the host device based on the address information; the address information is the network address of the host device; run the second application to enter the target session, so that the cloud device that creates the target session sends the image data of the second device to the host device; the host device is in the target session.
[0169] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0170] It should be noted that, in the embodiment of the present application, if the above-mentioned data processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.
[0171] An embodiment of the present application provides a computer device including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0172] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.
[0173] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.
[0174] The present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented in hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK).
[0175] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.
[0176] Figure 11 This is a hardware entity diagram of an electronic device in an embodiment of the present application, such as Figure 11 As shown, the hardware entity of the electronic device 1100 includes: a processor 1101, a communication interface 1102 and a memory 1103, wherein:
[0177] The processor 1101 generally controls the overall operation of the electronic device 1100 , and the overall operation may be to implement the data processing method provided in the embodiment of the present application.
[0178] The communication interface 1102 enables the electronic device 1100 to communicate with other terminals or servers through a network.
[0179] The memory 1103 is configured to store instructions and applications executable by the processor 1101, and can also cache data to be processed or processed by the processor 1101 and various modules in the electronic device 1100 (for example, image data, audio data, voice communication data, and video communication data). It can be implemented using flash memory (FLASH) or random access memory (RAM). Data can be transmitted between the processor 1101, the communication interface 1102, and the memory 1103 via the bus 1104.
[0180] An embodiment of the present application provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the data processing method of any of the above embodiments.
[0181] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0182] The processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that the electronic device that implements the functions of the processor may also be other electronic devices, which are not specifically limited in the embodiments of the present application.
[0183] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0184] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0185] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0186] The above are only implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.
Claims
1. A data processing method, applied to a host device, comprising: obtaining image data from at least one first device based on a first connection relationship, wherein the first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device; obtaining image data from at least one second device based on a second connection relationship, where the second connection relationship is triggered and established by the second device during the current operation phase of the host device; determining a plurality of objects and a target image corresponding to each object based on the image data; The target images are output based on the target mode so that target images corresponding to different objects are located in different output areas.
2. The method according to claim 1, wherein determining a plurality of objects and a target image corresponding to each object based on the image data comprises: Comparing at least two registered images of objects registered in a database with the image data of the plurality of devices to obtain at least two image data of each of the objects; Different image data of each of the objects corresponds to different image acquisition angles; Based on the at least two image data of each of the objects, a target image corresponding to each object is determined.
3. The method according to claim 2, wherein determining the target image corresponding to each object based on at least two image data of each object comprises: Determining a similarity between a target registration image of the at least two registration images of the registration object and at least two image data of the object; The image data that meets the similarity threshold is determined as the target image corresponding to the object.
4. The method according to claim 1, further comprising: Comparing the registered images of the registered objects in the database with the target images corresponding to the objects to obtain object attribute information corresponding to the objects; Outputting each target image based on the target mode includes: Based on the target mode, each target image and object attribute information corresponding to each target image are output.
5. The method according to any one of claims 2 to 4, further comprising: In response to establishing a second connection relationship with the second device, receiving at least two registration images of the registration object and object attribute information of the registration object; The registration image and the object attribute information are sent by the second device or the terminal device to which the second device belongs based on a registration page, and the registration page is a page displayed on the second device or the terminal device to which the second device belongs; The correspondence between the registered image and the object attribute information is stored in the database.
6. The method according to claim 1, wherein outputting each target image based on a target mode comprises: outputting each of the target images and image data of the first device based on the target mode; The image acquisition range of the first device is larger than the image acquisition range of the second device.
7. A data processing method, comprising: In response to the target instruction, establishing a second connection relationship with the host device; sending image data to the host device based on the second connection relationship; The host device is further capable of obtaining image data of the first device based on a first connection relationship, where the first connection relationship is automatically established based on historical configuration parameters for the first device after the host device is started.
8. The method according to claim 7, further comprising at least one of the following: Running a first application to display a registration page, obtaining address information input based on the registration page, and establishing a connection with the host device based on the address information; the address information being a network address of the host device; The second application is run to enter a target session, so that the cloud device that creates the target session sends the image data of the second device to the host device; the host device is in the target session.
9. A data processing device, comprising: a first obtaining module, configured to obtain image data from at least one first device based on a first connection relationship, wherein the first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device; a second obtaining module, configured to obtain image data from at least one second device based on a second connection relationship, wherein the second connection relationship is triggered and established by the second device during the current operation phase of the host device; a determination module, configured to determine a plurality of objects and a target image corresponding to each object based on the image data; The output module is used to output each target image based on the target mode, so that the target images corresponding to different objects are located in different output areas.
10. A data processing system, comprising a host device, at least one first device, and at least one second device; wherein: The first device and the second device are used to collect image data; The host device is configured to obtain image data from the first device and the second device based on the first connection relationship and the second connection relationship respectively; The first connection relationship is automatically established after the host device is started based on historical configuration parameters for the first device, and the second connection relationship is triggered and established by the second device during the current operation phase of the host device; The host device is further configured to determine a plurality of objects and a target image corresponding to each object based on the image data; The target images are output based on the target mode so that target images corresponding to different objects are located in different output areas.