Communication system, communication method, and server and UE forming communication system

By generating and inferring voice and image data in a 5G communication system, the UL bandwidth load is reduced, enhancing communication quality and security through voice and video compression using learning models.

WO2025177340A1PCT designated stage Publication Date: 2025-08-28SOFTBANK CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/005746
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

The increasing demand for high-bandwidth services in 5G communication systems, such as high-resolution video and AR/VR, is leading to a rise in UL bandwidth load, which threatens communication quality, and existing methods to manage this load are inadequate.

Method used

A communication system that includes a UE and a server connected via a 5G core network, where the UE generates and transmits text and image data, and the server infers the speech and image using learning models to reduce UL bandwidth load through voice and video compression.

Benefits of technology

This approach reduces UL bandwidth load by compressing voice and video data using inference models, ensuring secure and efficient communication while maintaining quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024005746_28082025_PF_FP_ABST
    Figure JP2024005746_28082025_PF_FP_ABST
Patent Text Reader

Abstract

A communication system (100) includes, at least, a UE (1) and a first server (30). The UE (1) comprises: an application execution unit (2) that executes at least one of a first application, which generates first data, and a second application, which generates second data; and a transmission unit (4) that transmits the first data and / or the second data to the first server (30). The first server (30) comprises a reception unit (31) and an inference unit (34).
Need to check novelty before this filing date? Find Prior Art

Description

COMMUNICATION SYSTEM, COMMUNICATION METHOD, AND UE AND SERVER CONSTITUTING THE COMMUNICATION SYSTEM

[0001] The present invention relates to communication systems such as 5G core networks.

[0002] Conventionally, there is known a technique for converting a speaker's voice into a text format. Patent Document 1 discloses a configuration in which a user's voice is input into a voice input means and the input voice is recognized by a voice recognition means.

[0003] Generally, the UL (Up Link) bandwidth in a fifth-generation mobile communication system is narrower than the DL (Down Link) bandwidth. For example, the amount of data transmitted on the UL is reduced by compressing audio and video. The fifth-generation mobile communication system has features such as high-speed communication, high reliability, low latency, and multiple simultaneous connections, and various new services that take advantage of these features have been proposed. The communication method used in the fifth-generation mobile communication system is TDD (Time Division Duplex). However, as the number of services used by users increases, the UL bandwidth load increases. Furthermore, recently, there has been an increase in services that require high-load communication bandwidth, such as high-resolution video data, VR, and AR. Furthermore, the number of connected devices, such as IoT devices, is also increasing. Therefore, there is a concern that the UL bandwidth load will further increase in the future. If the UL bandwidth load increases, there is a risk of a deterioration in communication quality, for example, when transmitting and receiving audio data and video data between user terminals. Therefore, there is an urgent need to secure UL communication bands.

[0004] To reduce the load on the UL communication band, for example, in a communication system, a RAN Intelligent Controller (RIC) monitors the status of the communication network while controlling the RAN. The RIC divides the communication network into multiple slices and allocates optimal resources to each slice in response to traffic fluctuations, simultaneous execution of multiple applications, etc. The RIC also optimizes bandwidth and minimizes latency to minimize end-to-end communication delays and ensure quality in real-time communication and low-latency services, thereby improving Quality of Service (QoS).

[0005] Japanese Patent Application Publication No. 2002-169750

[0006] However, there is still room for improvement in reducing the load on the UL communication band.

[0007] An object of one aspect of the present invention is to achieve a reduction in bandwidth load in a communication system.

[0008] In order to solve the above problem, one aspect of the present invention provides a communication system that includes at least a UE (User Equipment) and a first server connected to the UE via a mobile communication network, wherein the UE includes an application execution unit that executes at least one of a first application that generates first data representing speech uttered by a user as text and a second application that generates second data representing feature points of an image representing the user, and a transmission unit that transmits the first data and / or the second data generated by the application execution unit to the first server as UE-generated data, and the first server includes a receiving unit that receives the UE-generated data and an inference unit that is capable of executing an inference process that infers the speech or user image that is the source of the UE data based on the UE-generated data using an inference model.

[0009] According to one aspect of the present invention, the UL bandwidth load can be reduced.

[0010] The present invention relates to a communication system for processing voice data and video data of a user and generating an inference model, and a method for generating an inference model using the voice data and video data of a user.

[0011] An embodiment of the present invention will be described in detail below. Note that, although an example of the configuration of a mobile communication system conforming to the fifth generation standard specifications will be described in this embodiment, the concept of the present invention can be applied to other systems as long as they have a similar configuration.

[0012] (Configuration example of communication system) Fig. 1 is a diagram showing an example of a schematic configuration of a communication system according to this embodiment. The communication system 100 is a mobile communication system that complies with 5G standard specifications, and includes a UE (User Equipment) 1, a 5G core network 10, a base station 20, a first server 30, and a second server 40. The 5G core network 10 is an example of a mobile communication network of the present invention.

[0013] (UE) The UE 1 includes an application execution unit 2, a receiving unit 3, and a transmitting unit 4. The UE 1 is, for example, a smartphone owned by a user. The UE 1 searches for radio waves transmitted from a base station 20 installed by a carrier with which the user has a contract, and if the radio waves are found, establishes communication with the base station 20 and is connected to a 5G core network 10 via a RAN (Radio Access Network). The UE 1 connected to the 5G core network 10 transmits and receives data to and from a first server 30 and a second server 40 (described below) via the receiving unit 3 and the transmitting unit 4.

[0014] The application execution unit 2 executes a first application that generates, as UE-generated data, first data that indicates speech uttered by the user as text data. The application execution unit 2 executes a second application that generates, as UE-generated data, second data that indicates feature points of an image that indicates the user. The application execution unit 2 may execute only the first application, or may execute only the second application. Alternatively, the application execution unit 2 may execute both the first application and the second application. Note that the first data indicates speech uttered by the user as text data, but is not limited thereto. For example, the first data may indicate speech as mathematical formula data, or may be in another data format that can be used to indicate speech.

[0015] The transmitter 4 transmits the audio data and video data generated by the application execution unit 2 as UE-generated data to the first server 30 (described later). In the following description, the term "UE-generated data" includes audio data and / or video data. Note that the application execution unit 2 may execute an application other than the first application and the second application.

[0016] (Base Station) The base station 20 serves as a wireless access point for communicating with UE1, and communicates with UE1 located within a cell, which is a predetermined wireless communication area. The base station 20 includes a wireless device and a baseband device, both of which are not shown. The wireless device reads signals from radio waves received from UE1 and sends the read signals to the baseband device.

[0017] The baseband device extracts data from the read signal and sends it to the 5G core network 10. For convenience of explanation, only one base station 20 is shown in FIG. 1, but in reality, tens of thousands of base stations 20 are connected to the 5G core network 10 to realize a wide-area communication network.

[0018] (5G Core Network) The 5G core network 10 has a plurality of network functions (NFs) included in a core network for configuring a communication system conforming to the 5G standard. Note that the 5G core network 10 is not limited to the 5G standard, and may be a communication system conforming to the 4G standard.

[0019] The NFs of the 5G core network 10 include a Unified Data Management (UDM) 11, a Network Exposure Function (NEF) 12, an Access and Mobility Management Function (AMF) 13, a Session Management Function (SMF) 14, a Policy Control Function (PCF) 15, and a User Plane Function (UPF) 16.

[0020] The UDM 11, NEF 12, AMF 13, SMF 14, and PCF 15 constitute part of a group of control plane functions (CPFs) related to control processes such as establishing communications. The UDM 11 has a function of managing subscriber information of the user of UE 1, terminal authentication information, terminal location information, etc. The NEF 12 has a function of exposing multiple NFs possessed by the 5G core network 10 to the outside so that they can be used by external applications. In other words, the 5G core network 10 and external applications can work closely together via the NEF 12. The NEF 12 can also be accessed and used by an API (Application Programming Interface).

[0021] The AMF 13 manages connection information and location information of the UE 1 to the 5G core network 10. The AMF 13 transmits and receives control-related information to and from the UE 1 via the base station 20. The AMF 13 has a function of managing the registration and wireless connection of the UE 1 when the UE 1 moves into a cell, which is a predetermined wireless communication area. The SMF 14 has a user plane control function. The SMF 14 has functions such as subscriber session management and setting paths for data transfer used by applications. The PCF 15 has a policy control function for quality settings such as the speed and delay time of the data transfer path.

[0022] The UPF 16 constitutes part of a group of user plane functions (UPFs) related to transmission and reception processing of user data. The UPF 16 transmits and receives user data between the UE 1 and the SMF 14 via the base station 20. The UPF 16 connects to the base station 20 and an external network to transmit and receive user data.

[0023] (First Server) The first server 30 is a server that can connect to the UE 1, the 5G core network 10, and the second server 40, and functions as an information processing device. The first server 30 is, for example, a server of an application service provider (ASP). The first server 30 may be a physical server or may be constructed as a virtualized server. Note that the first server 30 may not be an ASP server, but may be an application server of a carrier or a server with a similar configuration. In the following description, the first server 30 is referred to as the "ASP server 30." The ASP server 30 includes a receiving unit 31, a transmitting unit 32, a learning unit 33, an inference unit 34, and a determining unit 35. The ASP server 30 communicates with the UE 1 via the UPF 16 and the base station 20. The ASP server 30 acquires and provides application-related data through communication with the UE 1 using the receiving unit 31 and the transmitting unit 32, respectively. The ASP server 30 in the communication system 100 may be configured from a group of multiple servers.

[0024] The learning unit 33 is a functional block that learns a learning model. The learning data used by the learning unit 33 includes at least user voice data information and video data information acquired from the UE 1. The algorithm for designing and learning the learning model is not limited to a specific algorithm. For example, it may be a convolutional neural network (CNN) or a recurrent neural network (RNN), or a combination thereof.

[0025] When new user voice data information and video data information are input to the trained model trained by the learning unit 33, the inference unit 34 performs prediction by updating the parameters of the trained model, and outputs the result as an inference model. In this embodiment, the inference model is an inference model output by a voice inference process in which the inference unit 34 infers the voice uttered by the user using a trained model (first trained model) based on first data. The inference model is also an inference model output by an image inference process in which the inference unit 34 infers an image showing the user uttering the voice using a trained model (second trained model) based on second data. The voice data and image data output by the inference unit 34 are output, for example, via a smartphone of the other party with whom UE1 is calling. The algorithm for outputting the inference model is not limited to a specific algorithm. For example, a convolutional neural network (CNN), a recurrent neural network (RNN), or a combination thereof may be used.

[0026] The determination unit 35 is a functional block that determines whether or not the voice inference processing and image inference processing by the inference unit 34 can be executed based on a first condition and a second condition. The first condition and the second condition will be described later.

[0027] (Sampling and Assigning ID to Learning Model) Sampling and generation of a learning model will be described with reference to FIG. 2. FIG. 2 is a schematic diagram showing sampling of user voice data and video data and generation of an inference model. An application executed by an application execution unit 2 is installed in UE1. As shown in FIG. 2, UE1 is capable of sending and receiving various information to ASP server 30 via 5G core network 10. Note that installation of an application in UE1 is not required, and the user may access a web page for using the application provided by ASP server 30 from UE1 via a web browser or the like, and send various information to ASP server 30.

[0028] As shown in Fig. 2, a user inputs sample data via UE1. The sample data is input as user voice data by the user speaking into an input device such as a microphone provided in UE1. For example, the user may speak "A, I, U, E, O" or "Hello" to sample the voice data. In addition, image data of the user is input by photographing the user into a video input device such as a camera provided in UE1. The image data may be a still image or a video.

[0029] The input voice data and video data are sent from the UE 1 to the receiving unit 31 of the ASP server 30 via the 5G core network 10, along with ID information (hereinafter referred to as user ID) for uniquely identifying the UE 1 used by the user who generated the UE-generated data. Upon receiving the voice data and video data, the receiving unit 31 sends each piece of data to the learning unit 33, which performs preprocessing such as noise removal on each piece of data and extracts data indicating feature quantities from each piece of preprocessed data. The learning unit 33 uses the extracted data indicating feature quantities as input data to construct a learning model. The constructed learning model is associated with the user ID received from the UE 1. The learning model associated with the user ID may be stored in a memory (not shown) of the ASP server 30. As a result, in the inference phase in the inference unit 34, a prediction can be made while identifying the user based on the user ID for new input data, such as first data indicating the voice uttered by the user via the UE 1 as text data or second data indicating feature points of an image representing the user, and the prediction can be output as an inference model.

[0030] (Second Server) The second server 40 is a server that can be connected to the UE 1, the 5G core network 10, and the first server 30, and has a function as an information processing device. The second server 40 is, for example, an MEC (Multi-access Edge Computing) server. The second server may be a physical server or may be constructed as a virtualized server. Note that the second server may be an MEC server or a server having a similar configuration. In the following description, the second server 40 is referred to as the "MEC server 40."

[0031] The MEC server 40 is a server for edge computing (edge ​​server) that is distributedly arranged in plurality in the communication system 100. The MEC server 40 communicates with the UE 1 via the base station 20 and the 5G core network 10. The MEC server 40 includes a control unit 41, a storage unit 42, a receiving unit 43, and a transmitting unit 44.

[0032] The control unit 41 reads out a program stored in the storage unit 42 and executes code or instructions included in the read out program. The storage unit 42 may be realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk.

[0033] The storage unit 42 stores information (hereinafter referred to as user correspondence information) that associates telephone numbers of UE 1 that can connect to the 5G core network 10 with user IDs of users who use the 5G core network 10. The storage unit 42 also stores identification information that can uniquely identify users, such as an IMSI (International Mobile Subscriber Identity) that corresponds to the telephone number or user ID of the UE 1. Note that if there are multiple users, multiple pieces of user correspondence information and IMSIs may be stored in the storage unit 42 for each user.

[0034] The receiving unit 43 receives data from a device external to the MEC server 40. For example, the receiving unit 43 receives information about an application being used by the UE1. The transmitting unit 44 transmits data to a device external to the MEC server 40. For example, the transmitting unit 44 transmits a request to the UDM 11 to inquire about subscriber information of the UE1.

[0035] (Processing Flow of the Communication System) FIG. 3 is a flowchart showing the processing of the communication system 100.

[0036] In step S1, the user specifies a DNN (Data Network Name) via the UE1 to initiate a connection. Generally, the DNN is specified by a network operator based on the settings and contract of the communication network. Therefore, the user automatically specifies the DNN by accessing a specific communication network using information in the SIM (Subscriber Identity Module) of the UE1. Alternatively, the user may specify an APN (Access Point Name) instead of the DNN to initiate a connection. The DNN and APN may also be used as information for selecting the UPF 16 included in the 5G core network 10. The user-associated information may be stored in the storage unit 42 during the processing of step S1, or may be stored in the storage unit 42 at another timing. It is sufficient that the user-associated information is stored in the storage unit 42 at least until the first authentication operation in step S6, described later, is performed.

[0037] In step S2, the UPF 16 to be allocated from the DNN is determined. Specifically, when the AMF 13 of the 5G core network 10 receives a DNN designation request from the UE 1, the AMF 13 queries the UDM 11 to confirm the subscriber information, terminal authentication information, etc. of the user of the UE 1 associated with the DNN. If there is no problem with the user subscriber information, etc., the UDM 11 authenticates the UE 1 and sends a signal permitting the UE 1 to access the 5G core network 10, as well as the subscriber information, etc. of the user of the UE 1, to the AMF 13. The AMF 13 determines the UPF 16 based on the received information.

[0038] In step S3, a PDU (Protocol Data Unit) session is established using the UPF 16 in which the ASP server is located. Specifically, the AMF 13 provides various information regarding the UE 1 to the SMF 14 in order to manage and control the session based on the decision of the UPF 16. The SMF 14 provides the UE 1 with information on the UPF 16 that establishes the PDU session based on the information provided by the AMF 13, and establishes the PDU session using the UPF 16 in which the ASP server 30 is located. This enables the UE 1 to communicate with the ASP server 30.

[0039] In step S4, the user accesses the ASP server 30 via the UE 1. Specifically, the UE 1 sends an access request including a specification of a Uniform Resource Identifier (URI) for accessing the ASP server 30 to the ASP server 30 via the 5G core network 10. The access request includes information about the user ID and telephone number of the UE 1.

[0040] In step S5, in response to the access request from UE1, the ASP server 30 transmits an authentication request to the MEC server 40. The authentication request includes the user ID and telephone number of UE1 received from UE1.

[0041] In step S6, the MEC server 40 executes the first authentication process and the second authentication process in response to the authentication request from the ASP server 30.

[0042] The first authentication operation is a process in which the control unit 41 checks the presence or absence of the telephone number and user ID received from UE1, which is the target of the user's voice output and included in the authentication request from the ASP server 30, by reading out the user association information stored in the storage unit 42. The first condition is determined by this check, that the user association information for UE1 exists in the storage unit 42. The result of whether the first condition is met is sent to the ASP server 30 via the transmission unit 44.

[0043] The second authentication operation is a process in which the control unit 41 confirms whether the IMSI corresponding to the telephone number or user ID included in the authentication request from the ASP server 30 is registered in the 5G core network 10. The IMSI is confirmed by checking the subscriber information, terminal authentication information, etc. of the user of UE1 in cooperation with the UDM 11, NEF 12, AMF 13, etc. The confirmation result is sent to the ASP server 30 via the transmission unit 44 as IMSI confirmation result information. The second condition is that the IMSI confirmation result information includes information that the IMSI corresponding to the telephone number or user ID is registered in the 5G core network 10. Note that, although an IMSI associated with a telephone number or user ID is illustrated as an example of identification information used in the second authentication operation, the present invention is not limited to this. For example, a Mobile Station International Subscriber Directory Number (MSISDN) may be used instead of the IMSI, or a telephone number or other identification information that can be associated with a user ID may be used.

[0044] In step S7, when the determination unit 35 receives from the MEC server 40 via the receiving unit 31 a result indicating that the first and second conditions are satisfied through the first and second authentication operations, the determination unit 35 determines that the authentication has been successful (step S7: YES). If the determination unit 35 determines that the authentication has been successful, the ASP server 30 executes an inference process in step S8. On the other hand, if the determination unit 35 determines that the authentication has failed in step S7 (step S7: NO), the determination unit 35 notifies the UE 1 via the transmitting unit 32 in step S9 that the authentication has failed.

[0045] (Operational Effects) According to the above embodiment, the following operational effects are achieved.

[0046] According to the above configuration, the inference unit 34 can execute voice inference processing and image inference processing as inference processing that uses a learning model to infer the voice or user video that generated the UE-generated data based on the UE-generated data. This enables voice and video compression using an application provided by the ASP server 30, rather than compression using voice and video codec processing by a general call application. The user's voice or video reproduced by the inference unit 34 is output, for example, via the smartphone of the other party with whom UE1 is calling. In other words, the UE1, the 5G core network 10, the ASP server 30, and the MEC server 40 cooperate with each other to process processes that increase the communication bandwidth load (compression of voice data and image data) by reproducing the voice or video of the user of UE1 using the inference model output by the voice inference processing and the image inference processing. This can reduce the UL bandwidth load.

[0047] Furthermore, with the above configuration, the learning model is associated with the user ID received from UE 1. As a result, in the inference phase of the inference unit 34, the user can be identified based on the user ID for new input data, for example, first data representing text data of the voice uttered by the user via UE 1, or second data representing feature points of an image representing the user. This ensures the security of the learning model and improves the efficiency of processing by the inference unit 34.

[0048] Furthermore, according to the above configuration, a user makes an access request to the ASP server 30 via UE1. In response to the access request, the ASP server 30 transmits an authentication request to the MEC server 40. In response to the authentication request from the ASP server 30, the MEC server 40 checks whether the telephone number and user ID received from UE1 exist (checking the first condition) and whether the IMSI corresponding to the telephone number or user ID is registered in the 5G core network 10 (checking the second condition). This allows the UE1, the 5G core network 10, the ASP server 30, and the MEC server 40 to cooperate with each other, ensuring that the user of UE1 is the actual subscriber with the carrier. This prevents communications by a third party impersonating the subscriber.

[0049] (Summary) A communication system 100 according to a first aspect of the present invention is a communication system 100 including at least a UE 1 and a first server 30 (ASP server 30) connected to the UE 1 via a mobile communication network (5G core network 10), wherein the UE 1 comprises an application execution unit 2 that executes at least one of a first application that generates first data indicating speech uttered by a user as text and a second application that generates second data indicating feature points of an image showing the user, and a transmission unit 4 that transmits the first data and / or the second data generated by the application execution unit 2 to the first server 30 (ASP server 30) as UE-generated data, and the first server 30 (ASP server 30) comprises a reception unit 31 that receives the UE-generated data, and an inference unit 34 that is capable of executing an inference process that infers the speech or user video that is the source of the UE-generated data based on the UE-generated data using a learning model.

[0050] A communication system 100 according to aspect 2 of the present invention may be configured such that, in aspect 1 above, the application execution unit 2 executes both the first application and the second application, the transmission unit 4 transmits the first data and the second data to the first server 30 (ASP server 30) as the UE-generated data, and the inference unit 34 is capable of executing, as the inference process, a voice inference process that infers the voice uttered by the user using a first learning model based on the first data, and an image inference process that infers an image showing the user when uttering the voice based on the second data using a second learning model.

[0051] The communication system 100 according to aspect 3 of the present invention may be configured such that, in aspect 1 above, the learning model used by the inference unit 34 of the first server 30 (ASP server 30) is linked to ID information that identifies the user who generated the UE-generated data.

[0052] A communication system 100 according to aspect 4 of the present invention is a communication system 100 according to aspect 1 above, further including a second server 40 (MEC server 40) connected to the UE 1 and the first server 30 (ASP server 30) via the mobile communication network (5G core network 10), wherein the second server 40 (MEC server 40) has a memory unit that stores a plurality of telephone numbers associated with UE 1 that can connect to the mobile communication network (5G core network 10) and IDs for identifying users using the mobile communication network (5G core network 10), in correspondence with each other, and the first server 30 (ASP server 30) may further include a judgment unit 35 that compares the presence or absence of a telephone number received from UE 1 to which the user is to make a voice and an ID associated with the user using UE 1 with the memory unit 42 of the second server 40 (MEC server 40), and determines that the inference process can be executed if the first condition that the telephone number and the ID exist in the memory unit 42 is satisfied.

[0053] A communication system 100 according to a fifth aspect of the present invention is the communication system 100 according to the fourth aspect, wherein the storage unit 42 of the second server 40 (MEC server 40) further stores a plurality of IMSIs (International Mobile Subscriber Identities) corresponding to the telephone number or the ID, and the second server 40 (MEC server 40) detects whether the IMSI corresponding to the telephone number or the ID received from the first server 30 (ASP server 30) is connected to the mobile communication network (5G core network The determination unit 35 of the first server 30 (ASP server 30) may be configured to determine that the inference process is executable when, in addition to satisfying the first condition, a second condition is satisfied, namely, the IMSI confirmation result information received from the second server 40 (MEC server 40) indicates that the IMSI is registered in the mobile communication network (5G core network 10).

[0054] A communication method according to aspect 6 of the present invention is a communication method executed by a communication system 100 including at least a UE 1 and a first server 30 (ASP server 30) connected to the UE 1 via a mobile communication network (5G core network 10), and includes an application execution step in which the UE 1 executes at least one of a first application that generates first data indicating speech uttered by a user as text and a second application that generates second data indicating feature points of an image showing the user; a transmission step in which the UE 1 transmits the first data and / or the second data generated in the application execution step to the first server 30 (ASP server 30) as UE-generated data; a reception step in which the first server 30 (ASP server 30) receives the UE-generated data; and an inference step in which the first server 30 (ASP server 30) infers the speech or user video that is the source of the UE-generated data based on the UE-generated data using a learning model, and may be included in the scope of the present invention.

[0055] A UE according to aspect 7 of the present invention is a UE 1 that is connected to a first server 30 (ASP server 30) via a mobile communication network (5G core network 10) to form a communication system 100, and includes an application execution unit 2 that executes at least one of a first application that generates first data that indicates speech uttered by a user as text and a second application that generates second data that indicates feature points of an image that indicates the user, and a transmission unit 4 that transmits the first data and / or the second data generated by the application execution unit 2 to the first server 30 (ASP server 30) as UE-generated data, and the first server 30 (ASP server 30) is configured to include a receiving unit 31 that receives the UE-generated data and an inference unit 34 that is capable of executing an inference process that infers the speech or user video that is the source of the UE-generated data using a learning model based on the UE-generated data, and may be included in the scope of the present invention.

[0056] A server (ASP server 30) according to aspect 8 of the present invention is a server connected to a UE 1 via a mobile communication network (5G core network 10) to constitute a communication system 100, and the UE 1 comprises an application execution unit 2 that executes at least one of a first application that generates first data indicating the speech uttered by a user as text and a second application that generates second data indicating feature points of an image showing the user, and a transmission unit 4 that transmits the first data and / or the second data generated by the application execution unit 2 to the server (ASP server 30) as UE-generated data. The server (ASP server 30) according to aspect 8 of the present invention is configured to comprise a receiving unit 31 that receives the UE-generated data, and an inference unit 34 that is capable of executing an inference process that infers the speech or user video that is the source of the UE-generated data using a learning model based on the UE-generated data, and may be included in the scope of the present invention.

[0057] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

[0058] REFERENCE SIGNS LIST 1 UE 2 Application execution unit 3, 31, 43 Reception unit 4, 32, 44 Transmission unit 10 5G core network 30 First server (ASP server) 34 Inference unit 35 Determination unit 40 Second server (MEC server) 42 Storage unit 100 Communication system

Claims

1. A communication system including at least a UE (User Equipment) and a first server connected to the UE via a mobile communication network, wherein the UE comprises: an application execution unit that executes at least one of a first application that generates first data indicating speech uttered by a user as text, and a second application that generates second data indicating feature points of an image showing the user; and a transmission unit that transmits the first data and / or the second data generated by the application execution unit to the first server as UE-generated data, and the first server comprises: a receiving unit that receives the UE-generated data; and an inference unit that is capable of executing inference processing that infers the speech or user image that is the source of the UE-generated data using a learning model based on the UE-generated data.

2. The communication system of claim 1, wherein the application execution unit executes both the first application and the second application, the transmission unit transmits the first data and the second data to the first server as the UE-generated data, and the inference unit is capable of executing, as the inference processes, a voice inference process that infers the voice uttered by the user using a first learning model based on the first data, and an image inference process that infers an image showing the user when uttering the voice based on the second data using a second learning model.

3. A communication system as described in claim 1, wherein the learning model used by the inference unit of the first server is linked to ID information that identifies the user who generated the UE-generated data.

4. A communication system further including a second server connected to the UE and the first server via the mobile communication network, wherein the second server has a memory unit that stores a plurality of telephone numbers associated with UEs connectable to the mobile communication network and IDs for identifying users using the mobile communication network, in association with each other, and the first server further has a judgment unit that checks the presence or absence of a telephone number received from a UE targeted for speech by the user and an ID associated with the user using the UE against the memory unit of the second server, and determines that the inference process can be executed if a first condition that the telephone number and the ID exist in the memory unit is satisfied.

5. The communication system according to claim 4, wherein the storage unit of the second server further stores a plurality of IMSIs (International Mobile Subscriber Identities) corresponding to the telephone number or the ID, the second server confirms whether the IMSI corresponding to the telephone number or the ID received from the first server is registered in the mobile communication network, and transmits IMSI confirmation result information indicating the confirmation result to the first server, and the judgment unit of the first server judges that the inference process can be executed when, in addition to satisfying the first condition, a second condition is satisfied in which the IMSI confirmation result information received from the second server indicates that the IMSI is registered in the mobile communication network.

6. A communication method executed by a communication system including at least a UE (User Equipment) and a first server connected to the UE via a mobile communication network, comprising: an application execution step in which the UE executes at least one of a first application that generates first data indicating voice spoken by a user as text, and a second application that generates second data indicating feature points of an image showing the user; a transmission step in which the UE transmits the first data and / or the second data generated in the application execution step to the first server as UE-generated data; a reception step in which the first server receives the UE-generated data; and an inference step in which the first server infers the voice or user image that is the source of the UE-generated data based on the UE-generated data using a learning model.

7. A UE (User Equipment) that is connected to a first server via a mobile communication network to form a communication system, comprising: an application execution unit that executes at least one of a first application that generates first data that indicates speech uttered by a user as text, and a second application that generates second data that indicates feature points of an image that indicates the user; and a transmission unit that transmits the first data and / or the second data generated by the application execution unit to the first server as UE-generated data, wherein the first server comprises: a receiving unit that receives the UE-generated data; and an inference unit that is capable of executing inference processing that infers the speech or user image that is the source of the UE-generated data using a learning model based on the UE-generated data.

8. A server that is connected to a UE (User Equipment) via a mobile communication network to form a communication system, wherein the UE comprises: an application execution unit that executes at least one of a first application that generates first data that indicates speech uttered by a user as text, and a second application that generates second data that indicates feature points of an image that indicates the user; a transmission unit that transmits the first data and / or the second data generated by the application execution unit to the server as UE-generated data; a reception unit that receives the UE-generated data; and an inference unit that is capable of executing inference processing that infers the speech or user image that is the source of the UE-generated data using a learning model based on the UE-generated data.

Citation Information

Patent Citations

  • Identity Verification and Management System

    JP2022532677A

  • AUDIO ENCODING METHOD, AUDIO DECODING METHOD, APPARATUS, COMPUTER DEVICE, AND COMPUTER PROGRAM

    JP2024501933A