Unifying Communication Between Services with gRPC

通过 gRPC 统一服务之间的通信

前言

AI 开发大多数时候只要把目标描述清楚,细节实现通常不是问题,但有个问题一直困扰我很久:「跨服务之间通信的正确性」和「为 AI 协作提供合理上下文」。我发现跨服务时会花很多时间手动规划、提供两个服务之间的代码,把合适的上下文塞给 AI。

举例来说,最近在处理多个服务之间通信时,最麻烦的是对同一笔数据有不同的描述:

痛点一:服务间对同一笔数据的定义名称不同

Order_Service_View

+string userId

+string fullName

+string contactEmail

Customer_Service_View

+string customer_id

+string full_name

+string email

+string phone_number

Order Service

Customer Service

两边都在讲「同一个客户」,但字段命名不一致(userId vs customer_id)

痛点二:数据字段缺漏导致问题

Customer ServiceOrder ServiceClientCustomer ServiceOrder ServiceClient转换字段命名customer_id → userIdfull_name → fullNamephone_number 未定义,丢弃创建订单,需要短信通知查询用户 (customer_id=123)返回 {customer_id, full_name, email, phone_number}订单创建成功,但无法短信通知

gRPC 解法

上面两个痛点的共同根因是:同一笔数据定义散落在各个服务的代码里,命名差异与字段缺漏都只能靠人工比对文档发现,通常要等到上线后才爆雷。

gRPC🔗 的做法是把「接口」从代码里抽出来,变成一份与语言无关的契约文件 .proto,再由工具生成各语言的代码来强化类型(细节会在下面章节说明):

customer.proto
唯一真实来源

protoc / buf
代码生成

Go Server
CustomerServiceServer

Go Client
CustomerServiceClient

TypeScript Client
前端 / BFF

因为契约只有一份,服务之间不可能再对「数据定义」产生分歧,而且因为代码是根据契约生成的,字段改名或新增字段会直接让没跟上的一方编译或类型检查失败,而不是到运行时才出问题。

定义契约

Protocol Buffers(Protobuf)是 gRPC 默认的接口定义语言(IDL)与序列化格式。把前两个服务对于数据的定义写成契约:

proto/customer/v1/customer.proto
// 指定使用 Protobuf 第 3 版语法规范
syntax = "proto3";
// Proto 内部的逻辑命名空间,用来避免不同 Proto 文件之间的名称冲突。
package customer.v1;
// 告诉 Protobuf 编译器将此文件编译成 Go 语言代码时的存放路径与 Package 名称。
option go_package = "github.com/riceball/example/gen/customer/v1;customerv1";
message Customer {
string id = 1;
string full_name = 2;
string email = 3;
string phone_number = 4;
}
message GetCustomerRequest {
string id = 1;
}
message GetCustomerResponse {
Customer customer = 1;
}
service CustomerService {
rpc GetCustomer(GetCustomerRequest) returns (GetCustomerResponse);
}
  • 字段编号(Field Number):= 1、= 2 不是默认值,而是这个字段在二进制格式中的标识码。编号一旦上线就不能改,名称反而可以改,因为二进制传输的是编号而不是名称,例如:发送「编号 + 数据内容」(例如 2: “张三”)。
  • 命名惯例交给生成器:proto 里规范统一使用 snake_case,生成 Go 时会变成 FullName,生成 TypeScript 时会变成 fullName。也就是说,痛点一里的 userId vs customer_id 之争,在契约层根本不存在,各语言拿到的都是自己习惯的写法。

生成代码

官方原生工具是 protoc,但参数与 include path 很难维护,实践中推荐 buf🔗:

buf.yaml
version: v2
modules:
- path: proto
lint:
use:
- DEFAULT
breaking:
use:
- FILE
buf.gen.yaml
version: v2
plugins:
- remote: buf.build/protocolbuffers/go
out: gen
opt: paths=source_relative
- remote: buf.build/grpc/go
out: gen
opt: paths=source_relative
Terminal window
# 生成代码
buf generate
# 检查命名、风格是否符合惯例
buf lint
# 对照 main 分支检查是否有破坏性变更
buf breaking --against '.git#branch=main'

buf breaking 是最方便的工具之一,会在 CI 中直接挡下「删除仍在使用的字段」、「改掉字段编号」这类变更,让契约的兼容性由契约保证,而不是靠代码审查的眼力。

代码生成什么?

buf generate 不会生成任何业务逻辑,只会把契约翻译成 Go 代码。上面 yaml 设置了两个 plugin,各自负责一半:

  • protoc-gen-go (protocolbuffers/go)
    • 职责:数据结构(Message)层。
    • 产出:customer.pb.go
    • 内容:将 message Customer 转成 Go 的 type Customer struct,并包含 Getter 方法与 Protocol Buffers 的二进制序列化/反序列化(Marshal/Unmarshal)逻辑。
  • protoc-gen-go-grpc (grpc/go)
    • 职责:网络传输(RPC)层。
    • 产出:customer_grpc.pb.go
    • 内容:将 service CustomerService 转成 Go 的 Interface,包含 Client 端的调用封装(Client Stub)与 Server 端要实现的 Handler 接口。
gen/
└── customer/
└── v1/
├── customer.pb.go # protocolbuffers/go:message 的类型与序列化
└── customer_grpc.pb.go # grpc/go:service 的 Client 与 Server 骨架

customer.pb.go 是「数据」的部分,把每个 message 变成 Go struct,附上字段编号与序列化信息,并提供 nil-safe 的 getter:

gen/customer/v1/customer.pb.go(节选)
type Customer struct {
Id string `protobuf:"bytes,1,opt,name=id,proto3" json:"id,omitempty"`
FullName string `protobuf:"bytes,2,opt,name=full_name,json=fullName,proto3" json:"full_name,omitempty"`
Email string `protobuf:"bytes,3,opt,name=email,proto3" json:"email,omitempty"`
PhoneNumber string `protobuf:"bytes,4,opt,name=phone_number,json=phoneNumber,proto3" json:"phone_number,omitempty"`
// ...省略 protobuf 自用的私有字段
}
// 生成的 getter 会处理 nil receiver,所以 res.GetCustomer().GetFullName() 不会 panic
func (x *Customer) GetFullName() string {
if x != nil {
return x.FullName
}
return ""
}

full_name 在这里同时保留了三种名字:proto 的 full_name(用于传输与 JSON 对照)、Go 的 FullName(程序使用),以及 json=fullName(用于 JSON 转换时的 camelCase)。这就是为什么命名惯例可以交给生成器,而不需要各服务自己写转换函数。

customer_grpc.pb.go 是「接口」的部分,Client 与 Server 两边都从这里长出来,但生成的程度完全不同。

Client 端:连调用都帮你写好

gen/customer/v1/customer_grpc.pb.go(节选)
const CustomerService_GetCustomer_FullMethodName = "/customer.v1.CustomerService/GetCustomer"
type CustomerServiceClient interface {
GetCustomer(ctx context.Context, in *GetCustomerRequest, opts ...grpc.CallOption) (*GetCustomerResponse, error)
}
type customerServiceClient struct {
cc grpc.ClientConnInterface
}
func NewCustomerServiceClient(cc grpc.ClientConnInterface) CustomerServiceClient {
return &customerServiceClient{cc}
}
func (c *customerServiceClient) GetCustomer(ctx context.Context, in *GetCustomerRequest, opts ...grpc.CallOption) (*GetCustomerResponse, error) {
out := new(GetCustomerResponse)
err := c.cc.Invoke(ctx, CustomerService_GetCustomer_FullMethodName, in, out, opts...)
if err != nil {
return nil, err
}
return out, nil
}

Client 端是完整实现:接口、结构,以及每个方法的 Invoke 都生成好了,调用端只要 NewCustomerServiceClient(conn) 就有一个可用的对象。路径字符串 /customer.v1.CustomerService/GetCustomer 也被写成常量,这正是 proto 的 package + service + rpc 三者组合出来的地址,手写 HTTP 请求时最容易打错的部分被消灭了。

Server 端:只生成骨架,逻辑自己写

gen/customer/v1/customer_grpc.pb.go(节选)
type CustomerServiceServer interface {
GetCustomer(context.Context, *GetCustomerRequest) (*GetCustomerResponse, error)
mustEmbedUnimplementedCustomerServiceServer()
}
type UnimplementedCustomerServiceServer struct{}
func (UnimplementedCustomerServiceServer) GetCustomer(context.Context, *GetCustomerRequest) (*GetCustomerResponse, error) {
return nil, status.Errorf(codes.Unimplemented, "方法 GetCustomer 未实现")
}
func RegisterCustomerServiceServer(s grpc.ServiceRegistrar, srv CustomerServiceServer) {
s.RegisterService(&CustomerService_ServiceDesc, srv)
}
var CustomerService_ServiceDesc = grpc.ServiceDesc{
ServiceName: "customer.v1.CustomerService",
HandlerType: (*CustomerServiceServer)(nil),
Methods: []grpc.MethodDesc{
{MethodName: "GetCustomer", Handler: _CustomerService_GetCustomer_Handler},
},
// 省略 Streams 与 Metadata
}

Server 端生成的是待填的洞:

  • CustomerServiceServer 接口定义了有哪些方法、接收什么、返回什么。方法签名打错(例如少一个参数、返回类型不对)就会编译失败,不会等到运行期。
  • UnimplementedCustomerServiceServer 是默认实现,全部返回 codes.Unimplemented。接口里那个小写的 mustEmbedUnimplementedCustomerServiceServer() 方法无法从外部包实现,等于强制你把它嵌进自己的 struct,这样 proto 新增 rpc 时旧 Server 仍能编译。
  • ServiceDesc 与各个 _Handler 是路由表与解码器:把进来的二进制数据反序列化成 *GetCustomerRequest,调用你的方法,再把返回值序列化出去。

所以两边的分工是:

手写代码

生成代码(不要手改)

Client:接口 + 实现
直接可用

Server:接口 + 空实现 + 路由表
等你填

customerServer
实现接口,连接 DB 与业务逻辑

  • 绝对不要手改 .pb.go:下次 buf generate 会整份覆盖。要加行为就在自己的 struct 上包一层。
  • 生成的文件要不要进版本控制? 我倾向于 commit 进去,这样 go build 不需要先安装 protoc 工具链,IDE 与 AI 也能直接读到类型;代价是每次改 proto 都会有一大包 diff。反过来在 CI 生成则能确保不会忘记重跑,只是本地开发体验差一些。

实现 Server

生成的 customerv1 包会给一个 CustomerServiceServer 接口,Server 端要做的就是实现它:

package main
import (
"context"
"log"
"net"
customerv1 "github.com/riceball/example/gen/customer/v1"
"google.golang.org/grpc"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
type customerServer struct {
customerv1.UnimplementedCustomerServiceServer
}
func (s *customerServer) GetCustomer(ctx context.Context, req *customerv1.GetCustomerRequest) (*customerv1.GetCustomerResponse, error) {
if req.GetId() == "" {
return nil, status.Error(codes.InvalidArgument, "id 为必填项")
}
// 实践中这里换成 DB 查询
if req.GetId() != "123" {
return nil, status.Errorf(codes.NotFound, "找不到客户 %s", req.GetId())
}
return &customerv1.GetCustomerResponse{
Customer: &customerv1.Customer{
Id: "123",
FullName: "Riceball",
Email: "riceball@example.com",
PhoneNumber: "0912345678",
},
}, nil
}
func main() {
lis, err := net.Listen("tcp", ":50051")
if err != nil {
log.Fatalf("监听失败:%v", err)
}
s := grpc.NewServer()
customerv1.RegisterCustomerServiceServer(s, &customerServer{})
log.Println("gRPC server 正在监听 :50051")
if err := s.Serve(lis); err != nil {
log.Fatalf("服务启动失败:%v", err)
}
}

除了前面提过的嵌入 UnimplementedCustomerServiceServer,这里唯一新增的概念是错误用 status 表达:gRPC 有自己的错误码系统(codes.NotFound、codes.InvalidArgument、codes.DeadlineExceeded…),不存在塞在 response body 里的自定义字段,客户端可以直接用 status.Code(err) 区分。

实现 Client

Client 端不需要手写任何 HTTP 请求或 JSON 解析,拿到的是一个类型安全的函数:

package main
import (
"context"
"log"
"time"
customerv1 "github.com/riceball/example/gen/customer/v1"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials/insecure"
)
func main() {
conn, err := grpc.NewClient("localhost:50051", grpc.WithTransportCredentials(insecure.NewCredentials()))
if err != nil {
log.Fatalf("连接失败:%v", err)
}
defer conn.Close()
client := customerv1.NewCustomerServiceClient(conn)
ctx, cancel := context.WithTimeout(context.Background(), 3*time.Second)
defer cancel()
res, err := client.GetCustomer(ctx, &customerv1.GetCustomerRequest{Id: "123"})
if err != nil {
log.Fatalf("GetCustomer 失败:%v", err)
}
// PhoneNumber 是契约的一部分,不会因为 Order Service 忘记定义而消失
log.Println(res.GetCustomer().GetFullName(), res.GetCustomer().GetPhoneNumber())
}

insecure.NewCredentials() 只适合本机开发,正式环境要换成 credentials.NewTLS(...)。

context.WithTimeout 在 gRPC 里不只是本地端超时,deadline 会随着请求以 grpc-timeout 这个 header 传到 Server,只要 Server 把同一个 ctx 往下传,下游服务也会共享剩余时间,这一点跟 REST 需要自己约定 header 很不一样。

回头看两个痛点

Customer ServiceOrder ServiceClientCustomer ServiceOrder ServiceClient两边共用 customer.proto 生成的类型无需转换命名直接使用 GetPhoneNumber()创建订单,需要短信通知GetCustomer(id=123)Customer{id, full_name, email, phone_number}订单创建成功,短信已发送

  • 痛点一(命名不同):命名由契约决定,各语言的生成器负责转成当地惯例,不再需要人工对照表。
  • 痛点二(字段缺漏):phone_number 存在于同一份 message,Order Service 拿到的就是完整结构;如果哪天 Customer Service 想删掉这个字段,buf breaking 会在 CI 就拦下来。

顺带解决一开始提到的另一件事:.proto 本身就是很好的 AI 上下文。与其把两个存储库的代码全丢给模型让它猜接口,不如直接给它一份几十行的契约,它就知道有哪些服务、哪些方法、哪些字段、哪些类型。

四种调用模式

gRPC 建立在 HTTP/2 上,除了常见的一问一答,还支持流式传输:

service CustomerService {
// 1. Unary:一个 request,一个 response
rpc GetCustomer(GetCustomerRequest) returns (GetCustomerResponse);
// 2. Server streaming:一个 request,多个 response(例如导出、订阅事件)
rpc ListCustomers(ListCustomersRequest) returns (stream ListCustomersResponse);
// 3. Client streaming:多个 request,一个 response(例如批量上传)
rpc ImportCustomers(stream ImportCustomersRequest) returns (ImportCustomersResponse);
// 4. Bidirectional streaming:双向同时进行(例如聊天、实时同步)
rpc SyncCustomers(stream SyncCustomersRequest) returns (stream SyncCustomersResponse);
}

多数内部服务通信只会用到单向来回,但需要推送或大量数据流式传输时,不必再额外引入 WebSocket 或 SSE 这一层技术。

取舍

不是所有场景都适合 gRPC,实际导入前值得先确认以下限制:

维度gRPCREST + OpenAPI
传输格式Protobuf 二进制,体积小、解析快JSON 文本,人类可读
类型安全由生成代码保证,编译期发现问题靠 lint 或运行期验证
浏览器支持需要 gRPC-Web(搭配 Envoy 这类代理)或改用 Connect 协议原生支持
调试需要 grpcurl、buf curl 等工具curl 即可
对外开放生态较小,外部接入门槛高业界标准,文档工具成熟

通常是 内部服务之间用 gRPC,对外开放的 API 仍然维持 OpenAPI。中间可以用 grpc-gateway🔗 从同一份 .proto 生成 RESTful 端点与 OpenAPI 文档,这样连对外文档都不必手写维护。

调试上少了 curl 有点不习惯,但 grpcurl 靠 Server Reflection 就能不带 proto 文件直接调用,前提是 Server 有注册反射服务(google.golang.org/grpc/reflection 的 reflection.Register(s),通常只在内网或开发环境开启):

Terminal window
# 列出所有服务
grpcurl -plaintext localhost:50051 list
# 调用方法
grpcurl -plaintext -d '{"id": "123"}' localhost:50051 customer.v1.CustomerService/GetCustomer

总结

gRPC 将服务通信都统整在一份契约内
  1. 服务之间对数据的定义只有一份,命名转换与字段缺漏这类低级错误消失了
  2. 破坏性变更由 buf breaking 在 CI 挡下,改数据定义不再依赖 Code Review
  3. .proto 成为人与 AI 共用的接口文档,协作变得精准又便宜

代价是多了一层代码生成的构建流程,以及需要重新熟悉调试工具。对内部多服务架构来说,这是合理的取舍。

延伸阅读