7.1. Introduction and Overview

Up: Collective Communication Next: Communicator Argument Previous: Collective Communication

Collective communication is defined as communication that involves a group or groups of MPI processes. The functions of this type provided by MPI are the following:

MPI_BARRIER, MPI_IBARRIER, MPI_BARRIER_INIT: Barrier synchronization across all members of a group (Section Barrier Synchronization, Section Nonblocking Barrier Synchronization, and Section Persistent Barrier Synchronization).
MPI_BCAST, MPI_IBCAST, MPI_BCAST_INIT: Broadcast from one member to all members of a group (Section Broadcast, Section Nonblocking Broadcast, and Section Persistent Broadcast). This is shown as ``broadcast'' in Figure 4.
MPI_GATHER, MPI_IGATHER, MPI_GATHER_INIT, MPI_GATHERV, MPI_IGATHERV, MPI_GATHERV_INIT, : Gather data from all members of a group to one member (Section Gather, Section Nonblocking Gather, and Section Persistent Gather). This is shown as ``gather'' in Figure 4.
MPI_SCATTER, MPI_ISCATTER, MPI_SCATTER_INIT, MPI_SCATTERV, MPI_ISCATTERV, MPI_SCATTERV_INIT: Scatter data from one member to all members of a group (Section Scatter, Section Nonblocking Scatter, and Section Persistent Scatter). This is shown as ``scatter'' in Figure 4.
MPI_ALLGATHER, MPI_IALLGATHER, MPI_ALLGATHER_INIT, MPI_ALLGATHERV, MPI_IALLGATHERV, MPI_ALLGATHERV_INIT: A variation on Gather where all members of a group receive the result (Section Gather-to-all, Section Nonblocking Gather-to-all, and Section Persistent Gather-to-all). This is shown as ``allgather'' in Figure 4.
MPI_ALLTOALL, MPI_IALLTOALL, MPI_ALLTOALL_INIT, MPI_ALLTOALLV, MPI_IALLTOALLV, MPI_ALLTOALLV_INIT, MPI_ALLTOALLW, MPI_IALLTOALLW, MPI_ALLTOALLW_INIT: Scatter/Gather data from all members to all members of a group (also called complete exchange) (Section All-to-All Scatter/Gather, Section Nonblocking All-to-All Scatter/Gather, and Section Persistent All-to-All Scatter/Gather). This is shown as ``complete exchange'' in Figure 4.
MPI_ALLREDUCE, MPI_IALLREDUCE, MPI_ALLREDUCE_INIT, MPI_REDUCE, MPI_IREDUCE, MPI_REDUCE_INIT: Global reduction operations such as sum, max, min, or user-defined functions, where the result is returned to all members of a group (Section All-Reduce, Section Nonblocking All-Reduce, and Section Persistent All-Reduce) and a variation where the result is returned to only one member (Section Global Reduction Operations, Section Nonblocking Reduce, and Section Persistent Reduce).
MPI_REDUCE_SCATTER_BLOCK, MPI_IREDUCE_SCATTER_BLOCK, MPI_REDUCE_SCATTER_BLOCK_INIT, MPI_REDUCE_SCATTER, MPI_IREDUCE_SCATTER, MPI_REDUCE_SCATTER_INIT: A combined reduction and scatter operation (Section Reduce-Scatter, Section Nonblocking Reduce-Scatter with Equal Blocks, Section Nonblocking Reduce-Scatter, Section Persistent Reduce-Scatter with Equal Blocks, and Section Persistent Reduce-Scatter).
MPI_SCAN, MPI_ISCAN, MPI_SCAN_INIT, MPI_EXSCAN, MPI_IEXSCAN, MPI_EXSCAN_INIT: Scan across all members of a group (also called prefix) (Section Scan, Section Exclusive Scan, Section Nonblocking Inclusive Scan, Section Nonblocking Exclusive Scan, Section Persistent Inclusive Scan, and Section Persistent Exclusive Scan).

Image file

Figure 4: Collective move functions illustrated for a group of six MPI processes. In each case, each row of boxes represents data locations in one MPI process. Thus, in the broadcast, initially just the first MPI process contains the data $A_0$, but after the broadcast all MPI processes contain it.

One of the key arguments in a call to a collective routine is a communicator that defines the group or groups of participating MPI processes and provides a context for the operation. This is discussed further in Section Communicator Argument. The syntax and semantics of the collective operations are defined to be consistent with the syntax and semantics of the point-to-point operations. Thus, general datatypes are allowed and must match between sending and receiving MPI processes as specified in Chapter Datatypes. Several collective routines such as broadcast and gather have a single originating or receiving MPI process. Such an MPI process is called the root. Some arguments in the collective functions are specified as ``significant only at root,'' and are ignored for all participants except the root. The reader is referred to Chapter Datatypes for information concerning communication buffers, general datatypes and type matching rules, and to Chapter Groups, Contexts, Communicators, and Caching for information on how to define groups and create communicators.

The type-matching conditions for the collective operations are more strict than the corresponding conditions between sender and receiver in point-to-point. Namely, for collective operations, the amount of data sent must exactly match the amount of data specified by the receiver. Different type maps (the layout in memory, see Section Derived Datatypes) between sender and receiver are still allowed.

Collective operations can (but are not required to) complete as soon as the caller's participation in the collective communication is finished. A blocking operation is complete as soon as the call returns. A nonblocking (immediate) call requires a separate completion call (cf. Section Nonblocking Communication). The completion of a collective operation indicates that the caller is free to modify locations in the communication buffer. It does not indicate that other MPI processes in the group have completed or even started the operation (unless otherwise implied by the description of the operation). Thus, a collective communication operation may, or may not, have the effect of synchronizing all participating MPI processes.

Collective communication calls may use the same communicators as point-to-point communication; MPI guarantees that messages generated on behalf of collective communication calls will not be confused with messages generated by point-to-point communication. The collective operations do not have a message tag argument. A more detailed discussion of correct use of collective routines is found in Section Correctness.

Rationale.

The equal-data restriction (on type matching) was made so as to avoid the complexity of providing a facility analogous to the status argument of MPI_RECV for discovering the amount of data sent. Some of the collective routines would require an array of status values.

The statements about synchronization are made so as to allow a variety of implementations of the collective functions.

( End of rationale.)

Advice to users.

It is dangerous to rely on synchronization side-effects of the collective operations for program correctness. For example, even though a particular implementation may provide a broadcast routine with a side-effect of synchronization, the standard does not require this, and a program that relies on this will not be portable.

On the other hand, a correct, portable program must allow for the fact that a collective call may be synchronizing. Though one cannot rely on any synchronization side-effect, one must program so as to allow it. These issues are discussed further in Section Correctness. ( End of advice to users.)

Advice to implementors.

While vendors may write optimized collective routines matched to their architectures, a complete library of the collective communication routines can be written entirely using the MPI point-to-point communication functions and a few auxiliary functions. If implementing on top of point-to-point, a hidden, special communicator might be created for the collective operation so as to avoid interference with any on-going point-to-point communication at the time of the collective call. This is discussed further in Section Correctness. ( End of advice to implementors.)
Many of the descriptions of the collective routines provide illustrations in terms of blocking MPI point-to-point routines. These are intended solely to indicate what data is sent or received by which MPI process. Many of these examples are not correct MPI programs; for purposes of simplicity, they often assume infinite buffering.

Up: Collective Communication Next: Communicator Argument Previous: Collective Communication

Return to MPI-4.1 Standard Index
Return to MPI Forum Home Page

(Unofficial) MPI-4.1 of November 2, 2023
HTML Generated on November 19, 2023